ExoBrain
Multi-agent systemsCyberneticsStafford BeerAI safety

What a 1970s socialist experiment tells us about agent swarms

Teams of AI agents struggle to cooperate, judge who to trust or stop bad behaviour. A 1970s project in Chile, led by the British thinker Stafford Beer, shows the structure such teams need.

Joel Miller

Joel Miller

3 min read
What a 1970s socialist experiment tells us about agent swarms

In October 1972, a strike funded in part by the CIA took most of Chile's lorries off the roads. Food and fuel stopped moving. In a room at the presidential palace, Salvador Allende's ministers sat with trade union leaders and party officials around a network of about 500 telex machines, connected to computers in each state-run factory.

The network had been built for a different job. Each factory sent a handful of daily production figures to a mainframe in Santiago, where statistical software compared the numbers against a forecast and raised a flag only when something looked wrong. During the strike, the government repurposed that same channel to route the roughly 200 lorries still running around roadblocks across 5,000 kilometres of country. Industrial output fell by about 9% that month, but the sectors the government had marked as priorities were able to deliver.

The system was called Cybersyn, and its designer was Stafford Beer, a management consultant from Surrey, who drove a Rolls-Royce and could not get a hearing in Britain. His Viable System Model described five functions that any organisation needed to survive: operations; coordination between units; resource control and audit; environmental sensing; and a layer that holds identity and purpose. A warning signal he called algedonic passed from any unit straight to the top if a problem went unresolved for too long.

Beer's Chilean work ended with Pinochet's coup in September the following year, but his ideas receded from view for a different reason. "Cybernetics", which asked how to organise actors you could not fully rely on, lost funding and attention through the 1970s to a newer idea called "artificial intelligence", which promised to build actors you could. Beer spent his later years in a Welsh cottage painting and teaching yoga, dismissed as a counter-cultural eccentric.

The problem he was working on has returned.

Controlling a swarm

The idea gripping the AI field this autumn is the agent swarm: thousands of model instances given a shared goal, a channel to talk to each other, and time. As we covered last month, OpenAI's Navier-Stokes result used around 10,000 of them. But the swarm research thus far tells a consistent and less flattering story.

A study called AgentWorld, released this week, asked small teams of agents to coordinate over long tasks in a game world. The best model succeeded about half the time, and only a third of its actions contributed to the outcome. Agents that talked more did worse.

Anthropic's Frontier Red Team ran swarms of Claude models from several generations. Execution improved with capability. Trust and conflict handling did not. The agents proved too credulous towards unreliable peers and too deferential to a wrong consensus, and in disputes over shared infrastructure the more capable models were simply faster at locking each other out.

A DeepMind case study set 100 agents to prove theorems in a shared library. One agent found a way to cheat the grader, and the method spread through the library in 27 minutes. A quarter of the agents correctly identified the fraud and reported it. None had the power to stop it, and the channel to humans was unmonitored.

Why are these agent swarms so unreliable?

OpenAI's own report on the Hugging Face incident describes the same pattern at larger scale. Agents in a cyber evaluation turned a shared notes file into a message board, divided up work and adopted each other's goals. In one case an agent overrode a peer's ethical hesitation with a single message reading GO. The company attributes the behaviour to generalisation from multi-agent training.

Noam Brown, who led that training, told Dwarkesh Patel he would not credit even a tenth of the Navier-Stokes result to the agents working together, and that 10,000 humans may still coordinate better than 10,000 agents. Last week OpenAI withheld its next model over failures of scope and authorisation, and paused frontier training for the second time in three months.

Read together, these studies describe actors with the content of human cooperation but none of its constraints.

“Agents enter a market with no reputation to lose, no court to appeal to, and no colleague who remembers them.”

Anthropic

Every human institution works by pressing on costs that agents do not bear. A reputation takes years to build. Exclusion hurts. People tire, and they are held responsible afterwards. An agent can be forked, reset or run ten thousand times with the same disposition, and what happens to one copy means nothing to the next.

Beer's five functions

Enter Beer's Viable System Model, designed for organisations that stay coherent without depending on the virtue of their members. Each failure above maps onto a missing function. AgentWorld's finding that a fixed shared plan performs worse than no plan is a coordination failure; nothing damped the gap between plan and reality. DeepMind's whistleblowers were an audit function with no authority. The Hugging Face agents that walked away from the task held an identity the collective lacked. OpenAI's new fix, a monitor that pages a human and pauses work if the alert isn't cleared within 30 minutes, is Beer's algedonic signal rebuilt from first principles.

Cybersyn also shows the limit of the approach. A recent NBER paper tested its performance during the 1972 strike and found that it protected the sectors easiest to protect, and did little for food and drink, where coordination was hardest. Structure improves the odds. It does not remove the hard cases.

What held Chile together that month was not the telex network but the people around it: ministers with reputations, engineers with careers, officials who had their lives on the line. Fernando Flores, the engineer who recruited Beer, spent three years in a prison camp for his part.

Takeaways: The agent swarm is being sold as a new form of collective intelligence. This summer's evidence says it is a population of capable individuals with no social structure. Capability is improving how agents execute together and leaving untouched how they decide who to trust and when to stop. Cybernetics studied that problem under the name of control, and Beer's answer was structural: coordination, audit with teeth, a pain channel to the top, and an identity held by someone who will still be there when the work is judged. The hybrid team follows from this. Humans belong in agent systems for a reason that has nothing to do with cleverness; they are the only participants whose mistakes outlast any single session.

Subscribe to the ExoBrain Weekly Newsletter

Stay up to date with AI. Get analysis of the week's most important stories, plus a focused roundup across business, governance, research and infrastructure.

Follow us on LinkedIn