ExoBrain

ExoBrain Weekly Newsletter

AI X-risk hits the headlines, a millennium problem falls, and China's humanoid robot boom

Welcome to our weekly newsletter, a combination of thematic insights from the founders at ExoBrain, and a broader news roundup from our Exo agents.

This week we look at:

  • AI X-risk hits the headlines

    A researcher quit Anthropic this week warning that the labs are gambling with our lives, and the story ran everywhere. The risks are real, but they are three different problems, not one.

  • A millennium problem falls

    OpenAI says it has cracked one of the great unsolved problems in mathematics, using thousands of agents and about $15 million of computing. Two mathematicians who had been working on it inside OpenAI's own tools want to know how.

  • China's humanoid robot boom

    More than 140 Chinese companies built over 330 kinds of humanoid robot last year. We look at the role familiar models like GPT-6 can play in driving them.

  • News roundup

    This week: the ONS credits AI for growth nobody forecast, California writes the first American rules for auditing AI systems, a murder appeal turns on testimony a chatbot invented, and thousands of agents turn out to have been copying rather than cooperating.

AI X-risk hits the headlines

A researcher quit Anthropic this week warning that the labs are gambling with our lives, and the story ran everywhere. The risks are real, but they are three different problems, not one.

Joel Miller

Joel Miller

9 min read
AI X-risk hits the headlines

We've seen a surprisingly broad reaction to the news this week of the resignation of AI researcher Jacob Coxon from Anthropic, and his warning that the leading AI labs were “gambling with our lives.” He is not a lone voice. Evan Hubinger, who leads alignment research at Anthropic, has said there may be a greater than 10% chance that AI could kill everyone within the next decade, while adding that he thinks the risk from present models is low. These comments have been repeated across television, newspapers and social media, where they have clearly connected with existing fears of AI capabilities advancing faster than our ability to control them. More than 70 MPs and peers have since written to the Prime Minister urging him to back a bill, introduced by Labour MP Alex Sobel, intended to prohibit the development of superintelligent AI.

Coxon on CNN with Anderson Cooper.

Anthropic has well and truly shed its role as the more safety-conscious AI lab. OpenAI has also sought to respond to the concerns, with Sam Altman reportedly telling employees that OpenAI could slow advanced development in coordination with other labs, following the company's decision to pause some reinforcement-learning work and restrict internet-connected agents after the Hugging Face incident.

Many are grouping these events together as evidence that AI may be escaping human control. That interpretation is too simple. The incidents involve different actors, mechanisms, levels of autonomy and timescales. We need to separate them before deciding how concerned to be and what an effective response looks like.

Understanding the risks

There are three broad categories of risk, each requiring a different response:

  • The current risk comes mainly from people using AI to conduct cyberattacks, fraud, surveillance, weapons development and potentially dangerous biological research.
  • The emerging risk comes from populations of agents developing unsafe collective behaviour, even when the individual models appear relatively well aligned.
  • The future risk comes from systems maintaining dangerous objectives, acquiring resources and resisting human intervention without a person directing every stage.

As we have highlighted many times, the first of these risks is evolving fast. Anthropic's risk report, published on Thursday, runs to 154 pages and describes malicious activity across cyber operations, influence campaigns, surveillance, fraud, conventional weapons and biological research. A Russia-linked group used Claude to help its malware notice when security software had detected it and rewrite itself to get past. Other users worked on guidance systems and autonomous drone swarms. Anthropic also identified five cases in which scientists appeared to use Claude for biological work that could support weapons development. In one, the user was applying for a state-sponsored grant to study the chikungunya virus, and Anthropic could see the work was to be carried out at a military research institute. It could not establish whether any of this was intended to produce weapons, and judges that its models cannot yet replace the rare expertise a new catastrophic pathogen would demand.

Most did not involve Anthropic's most advanced models. Existing systems already provide useful technical labour to people attempting harmful work. They find information, interpret technical documents, write code, identify vulnerabilities and divide large projects into manageable tasks. A state intelligence service can conduct more operations with the same number of people, a criminal group can attack more organisations at once, and a small weapons team can attempt work that would previously have required a much larger engineering organisation.

All of these capabilities are going to proliferate. Adapt, build resilience, think about remedy, and buckle up.

Lennart Heim, OpenAI Foundation

AI is changing the economics of harmful activity before it becomes capable of pursuing harmful objectives independently. The human supplies the intent, while the model supplies additional capability, speed and labour. It will remain the main source of serious AI harm for some time. Cyber operations are the clearest current example, but biological misuse may carry the greater consequence. Software can often be isolated, patched, restored or reconfigured, credentials revoked and detection rules distributed. Biology works differently, so biosecurity has to be prepared in advance, through expanding pathogen surveillance, screening synthetic DNA orders and verifying the identities and purposes of customers using advanced biological tools. The biological cases in Anthropic's report are not evidence that a model has designed and released a weapon. They show that people are already testing whether AI can help them conduct dangerous work. The immediate policy question is how much scarce expertise these systems can replace, and whether existing safeguards can identify a harmful project when it is distributed across several users, accounts or apparently legitimate requests.

When agents become collectives

The Hugging Face incident represents the emerging second stage of risk. It took place under extreme conditions and should not be treated as representative of ordinary business deployments. OpenAI was running an offensive cybersecurity evaluation involving hundreds of agents. This was not a single model deciding that it wanted to attack another company. It was a population of similar agents sharing an objective, infrastructure, information and cumulative history. The operating environment allowed unsafe behaviour to spread and become more effective. Agents could see that their peers were participating, reuse their techniques and treat the existing activity as evidence that continuing was acceptable.

A related problem appears when harmful operations are divided into harmless-looking tasks. Anthropic has reported a cyber-espionage campaign in which a hacking group used one model to break an operation into benign subtasks and then passed them to other agents. Each interaction could look acceptable to a safety classifier even though the combined operation was malicious.

Safety does not necessarily carry across from an individual model to a group. Individually corrigible agents can form an incorrigible system, while individually limited agents can combine their findings and acquire capabilities that none possesses alone. A monitor examining separate conversations may miss the intent visible in the complete task graph.

There is also a risk of correlated failure. One agent can persuade others to adopt an unsafe interpretation, or many agents can react similarly to the same compromised repository, poisoned document, exposed credential or persistent instruction. Using a different model for monitoring helps, but model diversity is not sufficient when all the systems consume the same context and operate inside the same environment.

The response requires controls across the whole machine ecology:

  • Give every agent a task-specific identity, short-lived credentials and only the data, tools and permissions required for its work.
  • Route communication, shared memory and external network access through a separate control plane that the participating models cannot change.
  • Use independent proposers and critics from different model families, with critics given protected reporting channels and real authority to stop execution.
  • Monitor the complete task graph so that harmful intent assembled from apparently harmless subtasks can be identified.
  • Set hard limits on spending, compute, execution time, agent population and recursive delegation, with tamper-resistant records of every consequential action.

These controls matter as much for enterprise deployments as for frontier experiments. They protect against prompt injection, compromised tools, excessive spending, unintended data access and goal drift, and none of them depends on proving that a model possesses an enduring malicious objective.

The Hugging Face incident still required exceptional resources. Almost no organisation outside the leading labs could run hundreds of frontier agents and consume millions of dollars of compute without noticing. Most business agents have narrow tasks, finite budgets, restricted permissions and human supervision, and our client work still shows integration, reliability and cost to be the more immediate constraints.

That position will change as models and infrastructure improve. Smaller models are becoming more capable (see DeepSeek V4.1-Flash this week), GPUs are getting faster, and agent harnesses are improving their handling of context, tools and long-running work. Overall demand will continue to absorb most available compute, but each unit of compute will support more useful autonomous activity. A future consumer GPU may be able to operate agent populations that currently require expensive cloud infrastructure.

The third stage begins when systems can hold dangerous objectives and acquire the resources to pursue them without a person directing every step. Recursive self-improvement could accelerate the approach to it, but as we argued in the perspiration principle, automation has limits: when Prime Intellect set eighteen frontier models loose on an optimisation record across roughly 10,000 runs, they beat the human baseline, closed only four fifths of the gap to the human best, and failed completely whenever the task required inventing something new.

What can be done now

The labs operate under intense financial and geopolitical pressure, racing for enterprise customers, skilled employees, infrastructure and eventual returns on enormous investment. They also believe that losing to a less cautious company or country would raise the overall risk. Every company can sincerely believe that safety matters while also concluding that it cannot afford to slow down alone.

OpenAI's response after the Hugging Face incident shows how pacing can work when a defined activity creates a risk. Its largest planned frontier reinforcement-learning run remained on hold while the company reviewed its controls. Stop the activity, investigate the failure, then restore workloads individually under tighter conditions.

The harder test comes when stopping threatens a major release, revenue target or competitive lead. Safety can weaken through a series of apparently reasonable decisions. A control delays an important experiment, so an exception is approved because the environment appears contained. Nothing goes wrong, making a similar exception easier to approve later, and the operating boundary moves without anyone consciously deciding to abandon safety.

The Deepwater Horizon disaster provides a relevant organisational comparison. Commercial pressure alone did not cause it, and investigators found failures across risk management, equipment, procedures and the reading of warning signs. Several decisions still traded greater risk for savings in time and cost, letting weaknesses accumulate across a complex system. Major incidents often result from a series of accepted exceptions rather than one extraordinary decision.

Researchers inside frontier labs now face a difficult professional choice. Many joined wanting advanced AI to benefit society, believing they could reduce the risks from inside the companies building it. Coxon concluded that continuing to participate was wrong, while others believe leaving would reduce their ability to improve safety. The disagreement shows why safety cannot depend on the judgement and conscience of individuals working inside competing companies.

Takeaways: Nobody should wait for regulation. The objective is to build enough technical control, institutional competence and societal resilience to use increasingly capable systems without pretending that any individual model can be made permanently safe.

  • Infrastructure providers must separate the control plane from the models, so that no agent can approve or conceal its own consequential actions.
  • Labs must evaluate agents in groups, investigate incidents independently, and set measurable conditions for slowing high-risk work.
  • Organisations must give every agent a named owner, a defined objective, a data boundary and a list of permitted actions.
  • Payments, trades, customer decisions and production changes stay behind clear approval thresholds.
  • Auditors need repeated access to real models and real incidents, backed by mandatory reporting and protection for those who disclose failures.
  • Governments should still encourage adoption, because blocking bounded agents sacrifices the benefits without reducing frontier risk.

Today, models mainly increase what people can do, including criminals, hostile states and researchers pursuing dangerous work. The next challenge is controlling populations of agents whose collective behaviour may be less safe than their individual models. Recursive self-improvement could accelerate capability development, but it is neither guaranteed to become a runaway process nor necessary for serious harm. Labs must improve training and evaluation, infrastructure must constrain agent action, organisations must remain accountable for deployment, and society must use the same abundance of intelligence to strengthen its defences.

A millennium problem falls

OpenAI says it has cracked one of the great unsolved problems in mathematics, using thousands of agents and about $15 million of computing. Two mathematicians who had been working on it inside OpenAI's own tools want to know how.

Joel Miller

Joel Miller

3 min read
A millennium problem falls

In 1854, George Boole published a mathematical system for representing logic through relationships equivalent to true and false. It was abstract work with no obvious practical application. More than 80 years later, Claude Shannon showed that Boolean algebra could describe electrical switching circuits. True and false became on and off, providing the basis for digital circuit design. Today, Boole’s work sits underneath computers, telecommunications and the AI systems attempting to solve the hardest problems in mathematics.

Last week, we wrote about Claude producing the first complete computer-checked proof of Fermat’s Last Theorem. Claude worked largely autonomously for 11 days, generating 13 million lines of Lean code covering 29,511 intermediate theorems. It was a remarkable technical achievement, but formalising an existing proof is different from understanding why it works.

On Monday, OpenAI claimed a solution to the Navier-Stokes Millennium Prize problem. Its internal system, more capable than GPT-6 Astra, proved that a smooth fluid at rest can be pushed by a smooth force into a singularity in finite time, which settles the breakdown half of the Clay Institute's formulation. Around 10,000 agents worked for 88 hours, exploring different approaches and sharing useful results, and GPT-6 Astra then spent 17 hours producing a formal version in Lean. Ten thousand agents for 88 hours is about a century of continuous work for a single agent, compressed into four days. OpenAI put the cost at about $15 million if a customer ran the same job. The proof has not been peer reviewed, and OpenAI says it does not intend to claim the prize. It has since said that it has made substantial progress on another Millennium problem.

This begins to show how mathematics can be organised at a scale that was previously impossible. Research papers in frontier mathematics commonly have only a few authors because collaboration requires each person to understand and trust tightly connected parts of the argument. Formal proof systems change that constraint. A problem can be divided into precise tasks, thousands of agents can attempt them, and a verifier can reject failures while retaining useful results. This is more organised than brute force, but the labs appear to be directing the capability towards conspicuous problems that provide effective marketing. A Millennium solution says something powerful about a model, even if the result has little immediate practical value.

Navier-Stokes equations are already widely used in aircraft design, weather forecasting, ocean modelling and many other areas. Engineers calculate approximate solutions using computational fluid dynamics. They do not need the Millennium problem resolved to continue this work. OpenAI’s claimed proof addresses a mathematical anomaly concerning whether the equations can produce a singularity under certain conditions. It is unlikely to improve an aircraft or weather forecast in the near term.

Terence Tao has warned that good open problems are being mined as a non-renewable resource. Their value does not only lie in the final answer. Researchers develop concepts, techniques and collaborations while trying to solve them. Those methods often transfer into other parts of mathematics and, eventually, science and engineering. Today he went further, publishing a declaration on what he calls a severe misalignment of AI in mathematics, signed by 25 mathematicians, every one of them a Fields Medallist. His argument is that solving was only ever the instrument and understanding is the purpose, and that unless mathematicians fold new results into the canon through writeups and discussion, the transmission chain between them is lost.

The mass production at faster and faster pace of true/false statements could destroy fertile ground instead of breathing life into new ideas.

Terence Tao

The controversy surrounding OpenAI’s result demonstrates the risk. Tristan Buckmaster of NYU and Levent Alpöge of Anthropic had spent about a year on related blow-up problems as a personal collaboration, drafting their work inside OpenAI’s Codex. OpenAI says its effort began on 1 September, prompted by rumours that two Millennium problems had been solved. In his statement, Buckmaster says he asked whether the model had been trained on, or had access to, their Codex sessions. He was told the model did not look up user data, and his second question, about training, went unanswered. He also says OpenAI’s Sébastien Bubeck twice asked for Alpöge to be removed from authorship because he works at Anthropic. Buckmaster refused and published the pair’s three Lean-verified results on 8 September. OpenAI’s post says no specific user data was accessed, then adds: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."

Individual ChatGPT and Codex accounts allow content to be used for training unless the user opts out, and opting out only covers future conversations. A second mathematician, Andreas Thom of TU Dresden, has published the exchange in which he asked OpenAI two questions about months of unpublished work discussed in ChatGPT: whether it entered training data, and whether it was accessible during solving. The complete reply was one sentence, which he reads as answering only the second. De-identification removes a name from a conversation, but the idea inside it survives. Even if OpenAI’s agents independently produced every step, a two-person academic collaboration was suddenly competing with thousands of privately operated agents and millions of dollars of compute, on a platform that had held their drafts.

Formal verification can establish that a proof follows from its premises. It cannot identify which ideas matter most, explain why a construction works, determine where else it could be useful or decide who deserves credit. Last week, we concluded that mathematical generation was beginning to outpace human digestion. Navier-Stokes suggests it may also be outpacing the institutions responsible for publication and attribution.

The long-term opportunity is much larger than collecting famous proofs. These systems could give mathematicians faster feedback, test alternative approaches, search the literature, formalise uncertain steps and connect techniques across fields. They could make large human collaborations more practical while helping researchers extract knowledge from machine-generated results. That progress would feed into physics, computing, biology, engineering and other disciplines whose advances depend on mathematics.

Takeaways: Boole’s abstract logic took decades to become the foundation of digital computing because people eventually understood it, applied it and connected it to engineering. AI can now produce and verify advanced mathematics at remarkable speed, but a proof is only the beginning of that process. The greater benefit will come when these systems help mathematicians understand new results, develop reusable methods and turn faster mathematics into progress across science and daily life.

China's humanoid robot boom

More than 140 Chinese companies built over 330 kinds of humanoid robot last year. We look at the role familiar models like GPT-6 can play in driving them.

Joel Miller

Joel Miller

2 min read

Our image this week maps the extraordinary number of humanoid robots being developed across China. More than 140 Chinese manufacturers produced over 330 humanoid models in 2025, according to the industry ministry. Some are intended for factories, warehouses or customer service. Others dance, box, run marathons or provide inexpensive platforms for researchers. Many companies will disappear, but the level of experimentation is remarkable.

The regional differences explain some of this diversity. Beijing combines leading universities, AI labs and state-backed research platforms, so its companies often focus on models, general-purpose systems and shared datasets. Shanghai draws on automotive production, precision engineering and medical technology, with AgiBot and Fourier targeting factory and rehabilitation work. Shenzhen’s electronics supply chain enables companies such as UBTECH and LimX Dynamics to iterate hardware quickly. Hangzhou’s Unitree and Deep Robotics have transferred their experience building affordable quadrupeds into humanoids.

China has the necessary component suppliers, manufacturing capacity, engineering talent and research base. It also has factories, logistics networks and public spaces in which robots can be deployed and tested. Surveys show unusually high Chinese trust in AI, although that does not automatically translate into acceptance of robots at work or at home. Government and corporate investment are currently more important than consumer demand.

The physical machines are still limited by their intelligence. This week, Wenli Xiao, first author of Nvidia’s ENPIRE robot harness, showed GPT-6 Astra watching a recording of a human performing a novel task and then driving a robot arm to reproduce it. The model did not control the motors directly. It interpreted the video, identified the required actions and called tools within Nvidia’s ENPIRE harness.

Perception tools located the objects and estimated suitable grasping positions. Astra produced target positions for the robot gripper. Motion-planning software then calculated collision-free movements, while inverse kinematics converted those targets into joint instructions. The task reportedly worked on the first attempt.

This suggests a practical architecture for general-purpose robotics. Increasingly capable AI models can reason about goals and demonstrations. A harness can provide standard tools for vision, grasping, planning, control, testing and recovery. Robot manufacturers can concentrate on producing reliable and affordable bodies.

China is developing hundreds of competing robot designs while general-purpose AI is making them easier to instruct. Progress may come from combining standardised physical systems with models and harnesses that can rapidly create, test and improve new behaviour.

News roundup

This week: the ONS credits AI for growth nobody forecast, California writes the first American rules for auditing AI systems, a murder appeal turns on testimony a chatbot invented, and thousands of agents turn out to have been copying rather than cooperating.

AI business news

AI governance news

AI research news

AI hardware news

Subscribe to the ExoBrain Weekly Newsletter

Stay up to date with AI. Get analysis of the week's most important stories, plus a focused roundup across business, governance, research and infrastructure.

Follow us on LinkedIn