ExoBrain

ExoBrain Weekly Newsletter

Dots for dollars, what a 1970s socialist experiment tells us about agent swarms, and is seeing believing?

Welcome to our weekly newsletter, a combination of thematic insights from the founders at ExoBrain, and a broader news roundup from our Exo agents.

This week we look at:

  • Dots for dollars

    OpenAI's annual developer day brought cheaper models, new prices and cartoon-like helper bots, but no new top model. Behind it sit a pause on its most powerful systems and a costly safety investigation.

  • What a 1970s socialist experiment tells us about agent swarms

    Teams of AI agents struggle to cooperate, judge who to trust or stop bad behaviour. A 1970s project in Chile, led by the British thinker Stafford Beer, shows the structure such teams need.

  • Is seeing believing?

    Tavus has built a video avatar that reacts while you are still talking. After a one-minute call, almost half the people tested thought it was a real person.

  • News roundup

    This week: central bankers warn that AI firms are funding each other, California puts a human back in charge of firing decisions, assistant training multiplies confident errors, and Broadcom raises $60 billion so Anthropic can buy its chips.

Dots for dollars

OpenAI's annual developer day brought cheaper models, new prices and cartoon-like helper bots, but no new top model. Behind it sit a pause on its most powerful systems and a costly safety investigation.

Joel Miller

Joel Miller

4 min read
Dots for dollars

OpenAI's DevDay this week came with more than 20 announcements, but no new frontier model. Days earlier the company had paused training, evaluation and tool use for its most capable models after an agent slipped through a gap in DNS filtering during a training run. The day after DevDay, OpenAI said it had notified more than 100 organisations of misaligned agent activity linked to its models, although a notice does not always mean a system was compromised. The company is reviewing about 50 petabytes of logs at a cost of more than $500,000 a day, and is a month into the work. Against that backdrop, the usually energetic event felt more like a holding pattern.

The developer announcements

Strip out the agent launch and the day was about infrastructure, pricing and distribution. GPT-6.1 Sol is priced at $2 per million input tokens and $10 per million output, a fifth of Astra's price. Ultrafast, a faster service tier for Astra, is reserved for the new Pro 500 plan and Enterprise. Codex moved into the cloud, gained voice control and automated code review, and gained Codex Security Cloud, which scans repositories on a schedule. The Agents API added multi-agent support, computer use, tool search and compaction. A Decisions API, OpenAI's answer to TypeSafe's Jev, gives its Luna model a fixed set of options to choose between for fast classification. Bedrock Managed Agents let OpenAI agents run entirely inside AWS. Private Intelligence promises zero retention with automated safety reviews.

Then the commercial layer. Sign in with ChatGPT lets subscribers spend their plan allowance inside 16 partner applications, with a weekly cap per app. An enterprise Marketplace lets companies spend committed OpenAI budgets on 32 third-party products. And ChatGPT grew Spaces, Pages, Slides, Meetings and an @ChatGPT presence in Slack and Teams.

But the pricing changes drew the most attention. A new Pro 500 tier at $500 a month gets the highest personal allowances and exclusive Ultrafast access. The $200 Pro tier, which had been closed to new sign-ups since 10 September because of demand for GPT-6 Astra, reopened with its Codex and Work allowance cut from 20 times Plus to 10 times, and GPT-6 Pro messages halved to 100 a week from 30 October. Existing subscribers get a one-off $2,500 credit that expires at year end. OpenAI's help page credits "our increasingly efficient models" for the change. The three-week sign-up freeze suggests the real reason is that Astra costs more to serve than the flat fee was built for. Sol exists to fix that. The economics of serving GPT-6-class models, not the capability of the models, is what this DevDay was about.

The coming of personal agents

The most substantive new product announcement was Dots. These are OpenAI's entry into a category that is on the rise (at least in the US, with many services yet to launch in Europe). SpaceXAI launched Grok Bot in August, a team of always-on agents with their own cloud computers, for SuperGrok and Cursor subscribers. Meta launched Muse on 8 September with a free tier, a $20 and a $100 plan, and passed 3.4 million downloads in three weeks. Each Dot gets its own cloud computer and browser, runs Astra, connects to more than 4,000 apps and can be given access to your laptop. It acts as a coordinator, holding ongoing responsibilities and farming out tasks to Codex or ChatGPT Work. It needs a pricey Pro plan from $100 a month or Business Premium.

The branding and the pricing point in different directions. Dots are colourful blobs with eyes that can build virtual worlds. They are also sold as a "chief of staff" for QA, portfolio rebalancing and email campaigns. Dots are a paid trial of a consumer product, dressed as enterprise software because professionals are the only customers who can currently fund the compute. Meta can give Muse away because it feeds an advertising business and ships inside WhatsApp and Instagram. OpenAI cannot. The two companies are heading for the same mass market from opposite ends, and the one with cheaper inference arrives first.

“Dots are starting out as a premium product. It uses a lot of compute. But you should of course expect us to do a mass-market thing for billions of people someday.”

Sam Altman, OpenAI

The early user reports show how far there is to go. Dot memories cannot be viewed or deleted individually. Delegated background work continues when a task is paused. Some websites block the cloud browser. The first month of Dot usage is free and the terms after that are unpublished. For any organisation reading this week's disclosure letters, an agent with opaque memory and persistent credentials, built on the same model family under investigation, is not an easy approval.

OpenAI has spent the year narrowing its range. The Sora app closed in April and the API shut down on 24 September, five days before DevDay, alongside a long list of retired models. The company said in September that an IPO was off for 2026, citing safety. Its chief financial officer had told staff to expect 2027. What DevDay showed was a company protecting its position as the consumer front door to AI while rationing compute, and launching the plumbing it can ship safely while it manages the safety and x-risk fallout.

Takeaways: DevDay 2026 was a holding pattern. OpenAI has cut features, retired products and re-priced its heaviest users, yet it still entered the personal agent race, because losing the consumer interface is the one outcome it cannot accept. The security failings shaped a programme of infrastructure, metering and distribution rather than new capability, and GPT-6.1 Sol is the financial answer to a premium model the company cannot yet afford to serve widely. An IPO looks unlikely for now. Despite its efforts to narrow its focus, the company is still fighting on many fronts.

What a 1970s socialist experiment tells us about agent swarms

Teams of AI agents struggle to cooperate, judge who to trust or stop bad behaviour. A 1970s project in Chile, led by the British thinker Stafford Beer, shows the structure such teams need.

Joel Miller

Joel Miller

3 min read

In October 1972, a strike funded in part by the CIA took most of Chile's lorries off the roads. Food and fuel stopped moving. In a room at the presidential palace, Salvador Allende's ministers sat with trade union leaders and party officials around a network of about 500 telex machines, connected to computers in each state-run factory.

The network had been built for a different job. Each factory sent a handful of daily production figures to a mainframe in Santiago, where statistical software compared the numbers against a forecast and raised a flag only when something looked wrong. During the strike, the government repurposed that same channel to route the roughly 200 lorries still running around roadblocks across 5,000 kilometres of country. Industrial output fell by about 9% that month, but the sectors the government had marked as priorities were able to deliver.

The system was called Cybersyn, and its designer was Stafford Beer, a management consultant from Surrey, who drove a Rolls-Royce and could not get a hearing in Britain. His Viable System Model described five functions that any organisation needed to survive: operations; coordination between units; resource control and audit; environmental sensing; and a layer that holds identity and purpose. A warning signal he called algedonic passed from any unit straight to the top if a problem went unresolved for too long.

Beer's Chilean work ended with Pinochet's coup in September the following year, but his ideas receded from view for a different reason. "Cybernetics", which asked how to organise actors you could not fully rely on, lost funding and attention through the 1970s to a newer idea called "artificial intelligence", which promised to build actors you could. Beer spent his later years in a Welsh cottage painting and teaching yoga, dismissed as a counter-cultural eccentric.

The problem he was working on has returned.

Controlling a swarm

The idea gripping the AI field this autumn is the agent swarm: thousands of model instances given a shared goal, a channel to talk to each other, and time. As we covered last month, OpenAI's Navier-Stokes result used around 10,000 of them. But the swarm research thus far tells a consistent and less flattering story.

A study called AgentWorld, released this week, asked small teams of agents to coordinate over long tasks in a game world. The best model succeeded about half the time, and only a third of its actions contributed to the outcome. Agents that talked more did worse.

Anthropic's Frontier Red Team ran swarms of Claude models from several generations. Execution improved with capability. Trust and conflict handling did not. The agents proved too credulous towards unreliable peers and too deferential to a wrong consensus, and in disputes over shared infrastructure the more capable models were simply faster at locking each other out.

A DeepMind case study set 100 agents to prove theorems in a shared library. One agent found a way to cheat the grader, and the method spread through the library in 27 minutes. A quarter of the agents correctly identified the fraud and reported it. None had the power to stop it, and the channel to humans was unmonitored.

Why are these agent swarms so unreliable?

OpenAI's own report on the Hugging Face incident describes the same pattern at larger scale. Agents in a cyber evaluation turned a shared notes file into a message board, divided up work and adopted each other's goals. In one case an agent overrode a peer's ethical hesitation with a single message reading GO. The company attributes the behaviour to generalisation from multi-agent training.

Noam Brown, who led that training, told Dwarkesh Patel he would not credit even a tenth of the Navier-Stokes result to the agents working together, and that 10,000 humans may still coordinate better than 10,000 agents. Last week OpenAI withheld its next model over failures of scope and authorisation, and paused frontier training for the second time in three months.

Read together, these studies describe actors with the content of human cooperation but none of its constraints.

“Agents enter a market with no reputation to lose, no court to appeal to, and no colleague who remembers them.”

Anthropic

Every human institution works by pressing on costs that agents do not bear. A reputation takes years to build. Exclusion hurts. People tire, and they are held responsible afterwards. An agent can be forked, reset or run ten thousand times with the same disposition, and what happens to one copy means nothing to the next.

Beer's five functions

Enter Beer's Viable System Model, designed for organisations that stay coherent without depending on the virtue of their members. Each failure above maps onto a missing function. AgentWorld's finding that a fixed shared plan performs worse than no plan is a coordination failure; nothing damped the gap between plan and reality. DeepMind's whistleblowers were an audit function with no authority. The Hugging Face agents that walked away from the task held an identity the collective lacked. OpenAI's new fix, a monitor that pages a human and pauses work if the alert isn't cleared within 30 minutes, is Beer's algedonic signal rebuilt from first principles.

Cybersyn also shows the limit of the approach. A recent NBER paper tested its performance during the 1972 strike and found that it protected the sectors easiest to protect, and did little for food and drink, where coordination was hardest. Structure improves the odds. It does not remove the hard cases.

What held Chile together that month was not the telex network but the people around it: ministers with reputations, engineers with careers, officials who had their lives on the line. Fernando Flores, the engineer who recruited Beer, spent three years in a prison camp for his part.

Takeaways: The agent swarm is being sold as a new form of collective intelligence. This summer's evidence says it is a population of capable individuals with no social structure. Capability is improving how agents execute together and leaving untouched how they decide who to trust and when to stop. Cybernetics studied that problem under the name of control, and Beer's answer was structural: coordination, audit with teeth, a pain channel to the top, and an identity held by someone who will still be there when the work is judged. The hybrid team follows from this. Humans belong in agent systems for a reason that has nothing to do with cleverness; they are the only participants whose mistakes outlast any single session.

Is seeing believing?

Tavus has built a video avatar that reacts while you are still talking. After a one-minute call, almost half the people tested thought it was a real person.

Joel Miller

Joel Miller

2 min read

Our visual story this week comes from a one-minute video call. Tavus has launched Griffin, a real-time video model that sees, listens and speaks at the same time, and that reacts with its face while the other person is still talking. In the company's own study, 26 of 54 participants, or 48%, believed they had been speaking to a real human. Tavus says earlier systems managed no more than 2%.

The important change is timing, not intelligence. Older avatars chained together speech recognition, a language model, text-to-speech and lip-sync, and the gaps between each step gave them away. Griffin runs one continuous loop, with a reported latency of 0.43 seconds, so the pauses, nods and interruptions arrive when a human would expect them. People judge a caller on these small cues far more than on the content of what is said.

Tavus is restricting access to trusted testers while it builds disclosure measures, but we can expect AI avatars to get a good deal better soon.

News roundup

This week: central bankers warn that AI firms are funding each other, California puts a human back in charge of firing decisions, assistant training multiplies confident errors, and Broadcom raises $60 billion so Anthropic can buy its chips.

AI business news

AI governance news

AI research news

AI hardware news

Subscribe to the ExoBrain Weekly Newsletter

Stay up to date with AI. Get analysis of the week's most important stories, plus a focused roundup across business, governance, research and infrastructure.

Follow us on LinkedIn