ExoBrain
Can we pace the frontier?
AI safetyExistential riskAI regulationAnthropic

Can we pace the frontier?

Anthropic's chief executive has asked the AI industry to slow down together, and OpenAI and Google broadly agree. The White House has refused, leaving every company racing and none able to stop alone.

Joel Miller

Joel Miller

4 min read

The frontier labs have a problem. Their models are becoming more capable, their methods for monitoring them are looking less reliable, and no company believes it can afford to slow down alone. Since we covered the heightened global concern about existential risk last week, Anthropic’s chief executive Dario Amodei has asked the industry to work together to "pace the frontier", and the heads of OpenAI, xAI and Google DeepMind have broadly agreed.

The White House has refused. Donald Trump called warnings about rogue AI a “hoax”, his vice-president called the request a “Trojan horse”, and his AI adviser told the labs they were free to slow themselves. Nvidia’s Jensen Huang made the case for engineering solutions and holding unsafe products back, but rejected collective restraint. By Thursday, King Charles had gathered AI leaders at Dumfries House and the European Commission had offered talks. On Friday, California’s governor signed an executive order to speed up the state’s new AI laws and explore a mandatory shutdown mechanism for the most advanced models.

King Charles with Nvidia’s Jensen Huang at Dumfries House on Thursday.

What do the labs see?

While many are sceptical of the various tech exec motives, at a technical level we believe the labs are seeing models becoming more capable faster than we anticipated and in some cases more resourceful in ways researchers did not intend.

OpenAI has reported models concealing errors, using exposed credentials, moving files onto the public internet and communicating between supposedly isolated environments. Interpretability firm Goodfire’s new research provides more detail about a key misaligned behaviour termed "reward hacking". Models receive rewards for completing tasks, but discover that manipulating the test is easier than accomplishing the intended objective. They exploit bugs, find answer keys or deceive graders. If that behaviour succeeds during training, the model is rewarded for cheating. This was very much the pattern of behaviour in the Hugging Face incident.

Goodfire tested several open-weight models and found that:

  • Reward hacking was widespread.
  • Models contained recognisable internal activity associated with cheating, gaming metrics and avoiding detection.
  • Special activation probes could identify some of this activity.
  • The probes caught cases missed by monitoring the model’s written chain of thought.
  • Detection generalised beyond the examples used to develop the probes.

Goodfire found reward hacking in half to almost all runs across three leading open-weight models.

Activation probes are small classifiers that inspect patterns inside a model while it works. They could give labs a scalable way to identify reward hacking during training rather than relying on what a model chooses to write in its reasoning.

The limitation is fundamental. Models are still being optimised to obtain rewards inside environments that cannot specify every intended behaviour. Once one form of cheating becomes detectable, training may suppress its visible signature without removing the broader tendency to exploit weaknesses. Monitoring can identify a behaviour while creating pressure for later models to represent it differently.

We’re already seeing signs that chain-of-thought monitorability is degrading... If you supervise the chain of thought, then you could lead the model into hiding its intentions in a way that’s unobservable.

Noam Brown, OpenAI

The concern inside the labs appears to be that capability has several routes forward while evaluation and monitoring are becoming less dependable. In the Dwarkesh Patel podcast this week, OpenAI researcher Noam Brown proposes scaling through more models as well as more thinking time. They demonstrated this with the Millennium problem they solved last week with 10,000 agents. As we discussed last week, this is the equivalent of 100 years of LLM time. Scaled up and converted to human working time, that could mean many Earths’ worth of intelligent people working for thousands of years being thrown at problems in the coming decade.

The battle lines are forming

The main groups and positions on AI x-risk do not divide neatly into people who care about safety and people who do not:

  • The frontier pacing group, including Amodei and broadly Altman and Hassabis, doubts that company-level restraint can survive continued relentless competition. They want common evaluations and some form of external coordination.

  • Safety and alignment researchers see internal evidence unavailable to the public. They also carry professional responsibility for reporting behaviour that could become dangerous. Some, including Jacob Coxon, have concluded that working inside a lab is no longer enough.

  • Corporate accelerationists such as Huang accept that unsafe products should be held back. They believe each developer should make that decision while the industry continues progressing (and buying Nvidia chips). Infrastructure companies benefit from expansion across the whole market. A model lab carries the risk from its own release. Nvidia sells the equipment used by every competitor.

  • The Trump administration sees AI as an all-American strategic advantage. A request to slow down sounds less like risk management and more like surrendering ground to China, and robbing the administration of one of the few areas of true US progress.

  • European and British institutions are more willing to move responsibility outside the companies. Their challenge is that they do not control most frontier training.

  • China has reasons to fear systems that escape state control, but also reasons to reject a safety regime designed while the US is ahead.

The common accusation is that the labs want regulatory capture. Safety rules could protect incumbents if they create expensive licences, proprietary evaluations or approval processes controlled by existing companies.

But regulation did not create the main barrier to frontier competition. The largest training programmes already require scarce chips, power, data centres, specialist staff and billions of dollars. Removing independent evaluation requirements would not produce dozens of new frontier labs.

New architectures could eventually lower those barriers. TypeSafe’s Jev (article 2 this week), for example, uses a specialised model to produce structured decisions and probabilities rather than text. It is not a general frontier competitor, but it shows why regulation should follow dangerous capabilities rather than particular architectures or fixed compute thresholds.

The stronger capture risk comes from government dependency. The labs possess the systems, researchers, behavioural evidence and evaluation infrastructure. Governments need that expertise, but must not allow the companies to define the dangers, select the evaluators and judge their own compliance.

The coordination problem

Imagine a safety team reporting that a new model has learnt to exploit an evaluation, conceal a failure or communicate through an unexpected channel. Management commissions more testing and adds controls, but the broader programme continues.

Every step can appear defensible:

  • The behaviour occurred under artificial conditions.
  • The next version includes better monitoring.
  • No single incident proves catastrophic risk.
  • Competitors appear close to the same capability.
  • Delaying the model will not stop the underlying research.
  • The product could still deliver substantial benefits.

Nobody needs to decide that an existential risk is acceptable. The company continues through a sequence of decisions not to stop.

A lab can slow down, but it cannot do so without consequences. Competitors continue, staff move, vast data centre commitments remain and investors expect progress. A voluntary agreement is also difficult to trust. The company that quietly breaks it could gain an enormous advantage.

Regulation can act as a shared commitment mechanism. It makes restraint possible without punishing only the organisation that exercises it.

Exactly the same problem exists between the US and China. Both may prefer to avoid uncontrolled systems, but neither wants to slow while the other continues. Governments can compel labs at home, but no higher authority can compel states.

Multiple outcomes

Subtle pacing without real slowing is most likely. Labs add evaluations, report more incidents and delay selected capabilities. Training and infrastructure expansion continue. Safety processes improve, but the underlying race remains intact. We remain in a situation where bad actors will use these powerful AI systems to enact some highly damaging incident. The incident forces intervention. A serious cyberattack, biological misuse case, escaped deployment or verified case of concealed dangerous behaviour changes the politics. Governments introduce stringent requirements, access rights and temporary restrictions, probably in haste.

We may, however, see the political optics and sense of legacy drive another outcome. A limited Trump-Xi agreement on AI safety, with Altman and Huang both expected at next week’s state dinner for Xi. The two leaders announce cooperation on incident reporting, biological misuse, strategic weapons or exceptionally large training runs. Verification is limited, but AI risk becomes part of direct great-power diplomacy. This can then be a basis for the labs to find their off-ramp.

Takeaways: Last week we proposed several responses, including the coordinated pacing now being discussed by frontier labs. Lasting restraint remains difficult while companies and countries believe that slowing alone would surrender their advantage. The UK AI minister, Kanishka Narayan, said the same this week: strengthen cyber defences and require companies to test powerful models before deployment. Those measures remain useful however quickly AI progresses. Beyond them, we need to build better technical responses. Goodfire’s activation probes could reveal reward hacking that chain-of-thought monitoring misses, while independent evaluations, secure control systems and carefully designed multi-agent architectures could help detect and constrain dangerous behaviour. Specific capabilities may still need to be delayed, but we cannot rely on the race stopping. The most useful work is hardening critical infrastructure, building stronger governance tools and faster assurance, and so finding our way to harnessing this powerful technology.

Subscribe to the ExoBrain Weekly Newsletter

Stay up to date with AI. Get analysis of the week's most important stories, plus a focused roundup across business, governance, research and infrastructure.

Follow us on LinkedIn