
In 1854, George Boole published a mathematical system for representing logic through relationships equivalent to true and false. It was abstract work with no obvious practical application. More than 80 years later, Claude Shannon showed that Boolean algebra could describe electrical switching circuits. True and false became on and off, providing the basis for digital circuit design. Today, Boole’s work sits underneath computers, telecommunications and the AI systems attempting to solve the hardest problems in mathematics.
Last week, we wrote about Claude producing the first complete computer-checked proof of Fermat’s Last Theorem. Claude worked largely autonomously for 11 days, generating 13 million lines of Lean code covering 29,511 intermediate theorems. It was a remarkable technical achievement, but formalising an existing proof is different from understanding why it works.
On Monday, OpenAI claimed a solution to the Navier-Stokes Millennium Prize problem. Its internal system, more capable than GPT-6 Astra, proved that a smooth fluid at rest can be pushed by a smooth force into a singularity in finite time, which settles the breakdown half of the Clay Institute's formulation. Around 10,000 agents worked for 88 hours, exploring different approaches and sharing useful results, and GPT-6 Astra then spent 17 hours producing a formal version in Lean. Ten thousand agents for 88 hours is about a century of continuous work for a single agent, compressed into four days. OpenAI put the cost at about $15 million if a customer ran the same job. The proof has not been peer reviewed, and OpenAI says it does not intend to claim the prize. It has since said that it has made substantial progress on another Millennium problem.
This begins to show how mathematics can be organised at a scale that was previously impossible. Research papers in frontier mathematics commonly have only a few authors because collaboration requires each person to understand and trust tightly connected parts of the argument. Formal proof systems change that constraint. A problem can be divided into precise tasks, thousands of agents can attempt them, and a verifier can reject failures while retaining useful results. This is more organised than brute force, but the labs appear to be directing the capability towards conspicuous problems that provide effective marketing. A Millennium solution says something powerful about a model, even if the result has little immediate practical value.
Navier-Stokes equations are already widely used in aircraft design, weather forecasting, ocean modelling and many other areas. Engineers calculate approximate solutions using computational fluid dynamics. They do not need the Millennium problem resolved to continue this work. OpenAI’s claimed proof addresses a mathematical anomaly concerning whether the equations can produce a singularity under certain conditions. It is unlikely to improve an aircraft or weather forecast in the near term.
Terence Tao has warned that good open problems are being mined as a non-renewable resource. Their value does not only lie in the final answer. Researchers develop concepts, techniques and collaborations while trying to solve them. Those methods often transfer into other parts of mathematics and, eventually, science and engineering. Today he went further, publishing a declaration on what he calls a severe misalignment of AI in mathematics, signed by 25 mathematicians, every one of them a Fields Medallist. His argument is that solving was only ever the instrument and understanding is the purpose, and that unless mathematicians fold new results into the canon through writeups and discussion, the transmission chain between them is lost.
“The mass production at faster and faster pace of true/false statements could destroy fertile ground instead of breathing life into new ideas.”
The controversy surrounding OpenAI’s result demonstrates the risk. Tristan Buckmaster of NYU and Levent Alpöge of Anthropic had spent about a year on related blow-up problems as a personal collaboration, drafting their work inside OpenAI’s Codex. OpenAI says its effort began on 1 September, prompted by rumours that two Millennium problems had been solved. In his statement, Buckmaster says he asked whether the model had been trained on, or had access to, their Codex sessions. He was told the model did not look up user data, and his second question, about training, went unanswered. He also says OpenAI’s Sébastien Bubeck twice asked for Alpöge to be removed from authorship because he works at Anthropic. Buckmaster refused and published the pair’s three Lean-verified results on 8 September. OpenAI’s post says no specific user data was accessed, then adds: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
Individual ChatGPT and Codex accounts allow content to be used for training unless the user opts out, and opting out only covers future conversations. A second mathematician, Andreas Thom of TU Dresden, has published the exchange in which he asked OpenAI two questions about months of unpublished work discussed in ChatGPT: whether it entered training data, and whether it was accessible during solving. The complete reply was one sentence, which he reads as answering only the second. De-identification removes a name from a conversation, but the idea inside it survives. Even if OpenAI’s agents independently produced every step, a two-person academic collaboration was suddenly competing with thousands of privately operated agents and millions of dollars of compute, on a platform that had held their drafts.
Formal verification can establish that a proof follows from its premises. It cannot identify which ideas matter most, explain why a construction works, determine where else it could be useful or decide who deserves credit. Last week, we concluded that mathematical generation was beginning to outpace human digestion. Navier-Stokes suggests it may also be outpacing the institutions responsible for publication and attribution.
The long-term opportunity is much larger than collecting famous proofs. These systems could give mathematicians faster feedback, test alternative approaches, search the literature, formalise uncertain steps and connect techniques across fields. They could make large human collaborations more practical while helping researchers extract knowledge from machine-generated results. That progress would feed into physics, computing, biology, engineering and other disciplines whose advances depend on mathematics.
Takeaways: Boole’s abstract logic took decades to become the foundation of digital computing because people eventually understood it, applied it and connected it to engineering. AI can now produce and verify advanced mathematics at remarkable speed, but a proof is only the beginning of that process. The greater benefit will come when these systems help mathematicians understand new results, develop reusable methods and turn faster mathematics into progress across science and daily life.