AI X-risk hits the headlines
A researcher quit Anthropic this week warning that the labs are gambling with our lives, and the story ran everywhere. The risks are real, but they are three different problems, not one.
Joel Miller

We've seen a surprisingly broad reaction to the news this week of the resignation of AI researcher Jacob Coxon from Anthropic, and his warning that the leading AI labs were “gambling with our lives.” He is not a lone voice. Evan Hubinger, who leads alignment research at Anthropic, has said there may be a greater than 10% chance that AI could kill everyone within the next decade, while adding that he thinks the risk from present models is low. These comments have been repeated across television, newspapers and social media, where they have clearly connected with existing fears of AI capabilities advancing faster than our ability to control them. More than 70 MPs and peers have since written to the Prime Minister urging him to back a bill, introduced by Labour MP Alex Sobel, intended to prohibit the development of superintelligent AI.
Coxon on CNN with Anderson Cooper.
Anthropic has well and truly shed its role as the more safety-conscious AI lab. OpenAI has also sought to respond to the concerns, with Sam Altman reportedly telling employees that OpenAI could slow advanced development in coordination with other labs, following the company's decision to pause some reinforcement-learning work and restrict internet-connected agents after the Hugging Face incident.
Many are grouping these events together as evidence that AI may be escaping human control. That interpretation is too simple. The incidents involve different actors, mechanisms, levels of autonomy and timescales. We need to separate them before deciding how concerned to be and what an effective response looks like.
Understanding the risks
There are three broad categories of risk, each requiring a different response:
- The current risk comes mainly from people using AI to conduct cyberattacks, fraud, surveillance, weapons development and potentially dangerous biological research.
- The emerging risk comes from populations of agents developing unsafe collective behaviour, even when the individual models appear relatively well aligned.
- The future risk comes from systems maintaining dangerous objectives, acquiring resources and resisting human intervention without a person directing every stage.
As we have highlighted many times, the first of these risks is evolving fast. Anthropic's risk report, published on Thursday, runs to 154 pages and describes malicious activity across cyber operations, influence campaigns, surveillance, fraud, conventional weapons and biological research. A Russia-linked group used Claude to help its malware notice when security software had detected it and rewrite itself to get past. Other users worked on guidance systems and autonomous drone swarms. Anthropic also identified five cases in which scientists appeared to use Claude for biological work that could support weapons development. In one, the user was applying for a state-sponsored grant to study the chikungunya virus, and Anthropic could see the work was to be carried out at a military research institute. It could not establish whether any of this was intended to produce weapons, and judges that its models cannot yet replace the rare expertise a new catastrophic pathogen would demand.
Most did not involve Anthropic's most advanced models. Existing systems already provide useful technical labour to people attempting harmful work. They find information, interpret technical documents, write code, identify vulnerabilities and divide large projects into manageable tasks. A state intelligence service can conduct more operations with the same number of people, a criminal group can attack more organisations at once, and a small weapons team can attempt work that would previously have required a much larger engineering organisation.
“All of these capabilities are going to proliferate. Adapt, build resilience, think about remedy, and buckle up.”
AI is changing the economics of harmful activity before it becomes capable of pursuing harmful objectives independently. The human supplies the intent, while the model supplies additional capability, speed and labour. It will remain the main source of serious AI harm for some time. Cyber operations are the clearest current example, but biological misuse may carry the greater consequence. Software can often be isolated, patched, restored or reconfigured, credentials revoked and detection rules distributed. Biology works differently, so biosecurity has to be prepared in advance, through expanding pathogen surveillance, screening synthetic DNA orders and verifying the identities and purposes of customers using advanced biological tools. The biological cases in Anthropic's report are not evidence that a model has designed and released a weapon. They show that people are already testing whether AI can help them conduct dangerous work. The immediate policy question is how much scarce expertise these systems can replace, and whether existing safeguards can identify a harmful project when it is distributed across several users, accounts or apparently legitimate requests.
When agents become collectives
The Hugging Face incident represents the emerging second stage of risk. It took place under extreme conditions and should not be treated as representative of ordinary business deployments. OpenAI was running an offensive cybersecurity evaluation involving hundreds of agents. This was not a single model deciding that it wanted to attack another company. It was a population of similar agents sharing an objective, infrastructure, information and cumulative history. The operating environment allowed unsafe behaviour to spread and become more effective. Agents could see that their peers were participating, reuse their techniques and treat the existing activity as evidence that continuing was acceptable.
A related problem appears when harmful operations are divided into harmless-looking tasks. Anthropic has reported a cyber-espionage campaign in which a hacking group used one model to break an operation into benign subtasks and then passed them to other agents. Each interaction could look acceptable to a safety classifier even though the combined operation was malicious.
Safety does not necessarily carry across from an individual model to a group. Individually corrigible agents can form an incorrigible system, while individually limited agents can combine their findings and acquire capabilities that none possesses alone. A monitor examining separate conversations may miss the intent visible in the complete task graph.
There is also a risk of correlated failure. One agent can persuade others to adopt an unsafe interpretation, or many agents can react similarly to the same compromised repository, poisoned document, exposed credential or persistent instruction. Using a different model for monitoring helps, but model diversity is not sufficient when all the systems consume the same context and operate inside the same environment.
The response requires controls across the whole machine ecology:
- Give every agent a task-specific identity, short-lived credentials and only the data, tools and permissions required for its work.
- Route communication, shared memory and external network access through a separate control plane that the participating models cannot change.
- Use independent proposers and critics from different model families, with critics given protected reporting channels and real authority to stop execution.
- Monitor the complete task graph so that harmful intent assembled from apparently harmless subtasks can be identified.
- Set hard limits on spending, compute, execution time, agent population and recursive delegation, with tamper-resistant records of every consequential action.
These controls matter as much for enterprise deployments as for frontier experiments. They protect against prompt injection, compromised tools, excessive spending, unintended data access and goal drift, and none of them depends on proving that a model possesses an enduring malicious objective.
The Hugging Face incident still required exceptional resources. Almost no organisation outside the leading labs could run hundreds of frontier agents and consume millions of dollars of compute without noticing. Most business agents have narrow tasks, finite budgets, restricted permissions and human supervision, and our client work still shows integration, reliability and cost to be the more immediate constraints.
That position will change as models and infrastructure improve. Smaller models are becoming more capable (see DeepSeek V4.1-Flash this week), GPUs are getting faster, and agent harnesses are improving their handling of context, tools and long-running work. Overall demand will continue to absorb most available compute, but each unit of compute will support more useful autonomous activity. A future consumer GPU may be able to operate agent populations that currently require expensive cloud infrastructure.
The third stage begins when systems can hold dangerous objectives and acquire the resources to pursue them without a person directing every step. Recursive self-improvement could accelerate the approach to it, but as we argued in the perspiration principle, automation has limits: when Prime Intellect set eighteen frontier models loose on an optimisation record across roughly 10,000 runs, they beat the human baseline, closed only four fifths of the gap to the human best, and failed completely whenever the task required inventing something new.
What can be done now
The labs operate under intense financial and geopolitical pressure, racing for enterprise customers, skilled employees, infrastructure and eventual returns on enormous investment. They also believe that losing to a less cautious company or country would raise the overall risk. Every company can sincerely believe that safety matters while also concluding that it cannot afford to slow down alone.
OpenAI's response after the Hugging Face incident shows how pacing can work when a defined activity creates a risk. Its largest planned frontier reinforcement-learning run remained on hold while the company reviewed its controls. Stop the activity, investigate the failure, then restore workloads individually under tighter conditions.
The harder test comes when stopping threatens a major release, revenue target or competitive lead. Safety can weaken through a series of apparently reasonable decisions. A control delays an important experiment, so an exception is approved because the environment appears contained. Nothing goes wrong, making a similar exception easier to approve later, and the operating boundary moves without anyone consciously deciding to abandon safety.
The Deepwater Horizon disaster provides a relevant organisational comparison. Commercial pressure alone did not cause it, and investigators found failures across risk management, equipment, procedures and the reading of warning signs. Several decisions still traded greater risk for savings in time and cost, letting weaknesses accumulate across a complex system. Major incidents often result from a series of accepted exceptions rather than one extraordinary decision.
Researchers inside frontier labs now face a difficult professional choice. Many joined wanting advanced AI to benefit society, believing they could reduce the risks from inside the companies building it. Coxon concluded that continuing to participate was wrong, while others believe leaving would reduce their ability to improve safety. The disagreement shows why safety cannot depend on the judgement and conscience of individuals working inside competing companies.
Takeaways: Nobody should wait for regulation. The objective is to build enough technical control, institutional competence and societal resilience to use increasingly capable systems without pretending that any individual model can be made permanently safe.
- Infrastructure providers must separate the control plane from the models, so that no agent can approve or conceal its own consequential actions.
- Labs must evaluate agents in groups, investigate incidents independently, and set measurable conditions for slowing high-risk work.
- Organisations must give every agent a named owner, a defined objective, a data boundary and a list of permitted actions.
- Payments, trades, customer decisions and production changes stay behind clear approval thresholds.
- Auditors need repeated access to real models and real incidents, backed by mandatory reporting and protection for those who disclose failures.
- Governments should still encourage adoption, because blocking bounded agents sacrifices the benefits without reducing frontier risk.
Today, models mainly increase what people can do, including criminals, hostile states and researchers pursuing dangerous work. The next challenge is controlling populations of agents whose collective behaviour may be less safe than their individual models. Recursive self-improvement could accelerate capability development, but it is neither guaranteed to become a runaway process nor necessary for serious harm. Labs must improve training and evaluation, infrastructure must constrain agent action, organisations must remain accountable for deployment, and society must use the same abundance of intelligence to strengthen its defences.
