Skip to content
← All notes
17 min read

Navier–Stokes and the End of Mathematics as We Knew It

AI Mathematics Formal Verification Technology

A few days ago, I heard a rumour that Claude had solved the Navier–Stokes problem. My reaction was fairly simple: I will believe it when I see it.

Apparently OpenAI heard a version of the same rumour.

On 1 September, after hearing that two Millennium Prize problems might have been resolved, OpenAI began testing a new internal model against all of the remaining open Millennium problems and several other research questions. On 8 September, it published what it says is a solution to Navier–Stokes, together with an analytical proof and a complete formalisation in Lean.

That last part matters. There is an enormous difference between an AI producing a hundred pages of convincing-looking mathematics and an AI producing a proof which a theorem prover accepts all the way down to the final statement. This is still a proposed solution, and it still needs independent mathematical scrutiny. The Clay Mathematics Institute has not declared the problem solved. But this is already much harder to dismiss than the stream of supposed solutions that famous open problems attract every year.

If the formal statement has been encoded correctly and the Lean development survives independent checking, then something very significant happened this week.

I find that exciting. I also find it unexpectedly sad.

Eighty-eight hours

The timeline in OpenAI’s account is almost more striking than the result itself.

OpenAI says it began training the relevant internal model on 28 August. By 1 September its researchers had seen enough of a step change in mathematical performance to launch a large evaluation across the open Millennium Prize problems. Different groups of agents were given different formulations and approaches. The agents could use code and a cached version of the internet, and groups could communicate internally.

They also gave the system easier neighbouring problems. One was the three-dimensional Euler regularity problem, essentially the zero-viscosity relative of Navier–Stokes. Nearly 100 agents worked on that for about 50 hours and produced a disproof of global regularity for the unforced Euler equations.

That changed the allocation of compute. OpenAI’s article contains a sentence which I suspect will look more important with time:

“To do so, we shifted agents away from the other Millennium Problems and prompted these agents with the Euler resolution.”

The other Millennium problems were not being discussed as some distant benchmark for a future system. OpenAI had agents actively working on them, decided Navier–Stokes looked most promising, and moved the agents over.

The group which produced the Navier–Stokes result involved roughly 10,000 concurrent agents. OpenAI says they reached the solution on 5 September, about 88 hours after the first agents were launched. Formalisation and verification in Lean then took another 17 hours, using GPT-6 Astra.

For Navier–Stokes alone, the agents exchanged 2.7 million messages and generated approximately 130 billion output tokens. Across all the attempted research problems, the totals were 4.9 million messages and about 300 billion output tokens.

It is worth keeping the scale in mind. This was not somebody opening ChatGPT, pasting in the Clay problem and receiving a proof before lunch. It was an industrial-scale research run using thousands of coordinated agents and an unreleased model which OpenAI says is significantly more capable than GPT-6 Astra. The model is also, according to OpenAI, still being trained.

That makes the result less magical. I am not sure it makes it less important.

There is also a messy question of priority in the background. The rumour OpenAI heard was later connected to work by Levent Alpöge and Tristan Buckmaster, who had themselves been using frontier AI systems on related fluid problems. OpenAI says neither its researchers nor its agents saw their work before it became public, while also saying it cannot completely rule out de-identified product data having contributed to model improvement. That dispute matters for attribution and for the norms around researchers using commercial AI tools. It deserves proper treatment of its own. For the narrower question here, the central issue is whether the released mathematics is correct and genuinely proves the stated theorem.

What has actually been claimed

The Navier–Stokes equations describe the motion of an incompressible viscous fluid. In one standard form,

ut+(u)u=νΔup+f,u=0.\frac{\partial u}{\partial t} + (u \cdot \nabla)u = \nu \Delta u - \nabla p + f, \qquad \nabla \cdot u = 0.

Here uu is the fluid velocity, pp is pressure, ν>0\nu > 0 is viscosity and ff is an external force.

The famous open question is whether sufficiently smooth three-dimensional data can evolve into a singularity in finite time. Viscosity smooths the flow, while the nonlinear transport term can concentrate it. For decades, nobody had been able to prove that smooth solutions always remain smooth, or construct an allowed example in which they break down.

OpenAI claims the second route. Its construction begins with a smooth fluid at rest and applies a smooth external force. The resulting flow remains finite in energy but develops unbounded velocity in finite time. The same programme is formalised both on R3\mathbb{R}^3 and on the periodic three-torus.

One subtlety is worth making explicit because headlines will inevitably lose it: this is a forced Navier–Stokes construction. OpenAI is not claiming to have found finite-time blow-up for the unforced equation. The reason it can nevertheless resolve the Millennium problem is that the official Clay formulation includes breakdown alternatives labelled (C) and (D) in which smooth forcing is allowed. OpenAI says it proves both of those alternatives.

That distinction makes the result a little less intuitive than the phrase “Navier–Stokes blows up” suggests, but it does not make it irrelevant to the stated prize problem if the formalisation really matches Clay’s conditions.

What Lean changes

AI-generated mathematics has had an obvious trust problem. A language model can produce ten pages in flawless mathematical prose and hide a fatal gap in line eleven. The better the model becomes at mathematical style, the less useful surface plausibility is as evidence of correctness.

This is where Lean changes things.

OpenAI’s formalisation repository describes the Navier–Stokes development as a full formalisation of the main results. Its metadata reports zero sorrys in those results and lists only the ordinary logical axioms used throughout Lean and Mathlib: propositional extensionality, classical choice and quotient soundness. The project also includes a Comparator setup against a Navier–Stokes statement adapted from Google DeepMind’s independent Formal Conjectures project.

If those claims reproduce under independent builds, Lean has checked every formal inference needed to reach the encoded theorem. For a proof released only days after the run began, that is hard evidence to wave away.

It is not the end of the checking process.

Formal verification proves the theorem which was formalised. Humans still have to establish that the definitions and final statement faithfully represent the mathematical problem we intended to ask. In a result like this, that means checking the exact regularity classes, forcing assumptions, energy conditions, domains and quantifiers against the official Clay formulation. Reviewers should also inspect the project for any accidental escape hatch and reproduce the build with the stated toolchain. OpenAI itself currently labels the formalisation’s review status as “self-assessed”.

So I would not update a textbook tonight to say that Navier–Stokes is unconditionally settled. OpenAI has released a proposed solution with a complete machine-checked proof behind it. That puts it in a very different category from an AI manuscript that merely looks convincing, but it still needs independent review.

Clay has another reason not to pronounce immediately. Under the Millennium Prize rules, a proposed solution must appear in a qualifying publication, survive at least two years of examination, and gain general acceptance in the global mathematics community before Clay will consider it for the prize. OpenAI has also said that it does not intend to claim the prize.

The caveat is real. It should not become an excuse to pretend nothing happened.

The part I find sad

There is something slightly dishonest about describing all of this only as exciting.

The Millennium Prize problems were deliberately chosen as monuments to the mathematical frontier. Clay’s own description says that part of the point was to show the public that mathematics still had deep open territory and to recognise achievements of historical magnitude. These problems acquired a cultural status beyond their technical definitions. They were mountains.

When people imagined one of them falling, they imagined something like Perelman and the Poincaré conjecture: years of thought, an extraordinary individual or small group, new mathematical ideas, a proof which becomes part of the story of a human life. The difficulty of the problem and the rarity of the person capable of solving it were tied together.

A datacentre containing ten thousand copies of a model changes the emotional character of that story.

Part of the unease goes beyond mathematics. Intelligence has always sat very close to the centre of how humans explain what is special about us. We are not the strongest or the fastest animal, but we reason, abstract, invent theories and prove things. Mathematics is perhaps the purest version of that self-image: thought with almost nothing else in the loop. A machine becoming stronger than us never felt particularly existential. A machine beginning to outrun us at one of the activities we use as a benchmark for abstract thought does.

The theorem is no less true because a machine found it. The mathematics is not less beautiful. But if this result holds, one of the traditional measures of mathematical greatness has become detached from human cognitive limits. A problem can resist the world’s best mathematicians for almost a century and then, once the right system exists, collapse over the course of a long weekend.

That is a strange thing to watch happen to a subject whose heroes are largely defined by the problems they could solve.

Chess offers a partial analogy. Human chess did not disappear when engines became stronger than every human player. People still play, study and care deeply about it. But nobody now confuses “best chess player” with “best entity at chess”. The machine ceiling moved far beyond the human one.

Mathematics is different because research is not mainly a competitive performance. Producing new mathematics is the substance of the profession. If machines become better at discovering proofs, inventing constructions and checking them formally, then the change reaches further than having a superhuman opponent on a chessboard.

The optimistic answer is that mathematicians move up a level. They choose worthwhile questions, formulate definitions, steer agents, identify interesting structures, explain machine-generated proofs and decide which results matter. I think that is probably what happens first. It may be a very productive way to do mathematics.

I am less convinced that it is a permanent refuge.

The OpenAI system was already being run as a population of agents exploring different approaches, sharing intermediate results and being redirected when the Euler result made Navier–Stokes look more promising. Human researchers still made important decisions in that process, but “strategy” is clearly entering the space of things we are learning to automate. The safe description of a mathematician’s future role keeps moving upwards as the systems improve.

Correctness and human understanding may separate as well. Today, a major theorem is expected to have a proof which experts can in principle read and internalise. In a world of machine-generated formal mathematics, we could have a vast frontier of theorems whose Lean certificates are trustworthy but whose best human explanation arrives months or years later. Mathematics could become partly an interpretive discipline: not only discovering what is true, but trying to understand a body of already verified machine mathematics which has advanced beyond the pace at which people can absorb it.

That would still be mathematics. It would not feel quite like the mathematics we inherited.

Caltech is already running the experiment

The timing of the Caltech Mathathon is almost comical.

From 30 October to 1 November, Caltech plans to host what it describes as the first hackathon devoted to research-level mathematics. One hundred teams will receive access to frontier models and spend 40 hours trying to solve open problems or build new mathematical theories. The headline offer is $2 million or more in AI credits, not a $2 million cash prize, and participants will later defend their work before leading mathematicians.

The event’s own page asks the question directly: “What is the role of a mathematician when AI can solve conjectures faster?”

A month ago that would have sounded like a provocative premise for a hackathon. It now reads more like a job description problem.

The likely near-term answer is that mathematicians become much more like principal investigators directing fleets of mathematical workers. Taste becomes more valuable. So does the ability to ask a precise question, recognise a promising construction and explain why a formally correct result is significant. Formal verification also becomes central, because human peer review will not scale if machines can produce difficult proofs faster than people can read them.

There is a very good version of this future. Open problems which would have consumed decades of human effort can be explored in days. Fields can test conjectures at a speed which changes how theory develops. Failed approaches become cheap. Formal proof libraries grow rapidly. A mathematician with a good question gains something like a laboratory containing thousands of tireless research collaborators.

There is also a less romantic version. The scarce resource stops being the ability to prove the theorem. It becomes compute, access to the strongest models, and perhaps the judgement needed to decide what to ask them.

Do the other Millennium problems now fall in a month?

The stupid extrapolation is irresistible.

If Navier–Stokes took 88 hours and there would be five unsolved Millennium problems left after an accepted solution, then five more at the same rate would take 440 hours: a little over eighteen days.

That calculation is obviously not a forecast.

OpenAI itself gives us the reason not to treat 88 hours as a universal constant. It started by allocating agents across all of the open Millennium problems and then concentrated on Navier–Stokes because the nearby Euler result made that route look unusually promising. Selection bias is doing a lot of work in the headline number. We are observing the problem which looked tractable enough to receive ten thousand agents, not a random sample from the six.

Navier–Stokes may also be unusually susceptible to this style of attack. The successful route is constructive: find a particular smooth setup whose dynamics produce a singularity while satisfying a long list of analytic constraints. That gives a large search space of possible ansätze, cancellations and geometric constructions which can be explored in parallel and checked locally. The Euler problem provides a closely related stepping stone. Recent human work in the same area is another indication that this part of the frontier was ripe.

Some of the remaining problems have a different character.

A proof that PNPP \ne NP, if that is the truth, probably requires a general lower-bound argument capable of escaping several famous barriers which have already ruled out broad families of proof techniques. There is no obvious single counterexample object to search for and then verify.

The Riemann hypothesis is a global statement about every non-trivial zero of the zeta function. If it is false, one sufficiently exotic zero would settle it. If it is true, a proof has to control an infinite analytic structure in a way no known method can currently do.

The Hodge and Birch–Swinnerton-Dyer conjectures sit inside deep networks of algebraic geometry and arithmetic geometry where progress often comes from inventing new conceptual bridges rather than tuning an explicit construction. Yang–Mills and the mass gap asks for a mathematically rigorous quantum field theory with the required physical property; even stating all of the machinery at the right level of rigour is part of the difficulty.

None of those descriptions amount to an argument that AI cannot solve them. In fact, the uncomfortable lesson of the last few years is that saying a task “requires a genuinely new idea” is no longer much of a defence. Models have repeatedly acquired abilities nobody separately programmed into them, and agent systems give those abilities time, tools, memory and enormous parallel search.

The line from OpenAI’s article matters because it is so matter-of-fact. They shifted agents away from the other Millennium Problems. The others were already in the queue.

I do not think we should expect one to fall every 88 hours. I also no longer think it is sensible to assume that the remaining five belong safely to another generation.

AGI, if the label still means anything

Five days before the Navier–Stokes announcement, OpenAI launched GPT-6 Astra, calling it “a new generation of intelligence”. At the launch briefing, OpenAI president Greg Brockman went further: he said he personally thought the company might have reached AGI and ended with “Welcome to the AGI era”.

I would not place much weight on treating that as a clean scientific threshold. “AGI” has been an unstable term for years, and by now it is partly a marketing category. There is no agreed experiment which turns a model from non-general to general at a particular score. Astra’s 98% on FrontierMath Tier 4 and 99.9% on ARC-AGI-3 tell us at least as much about the approaching end of those benchmarks as they do about where some metaphysical boundary called AGI ought to sit.

Language models have made the boundary especially blurry. Nobody separately programmed a translation engine, a proof engine, a coding engine and a planning engine into them. A broad training process produced all of those abilities together, and post-training, tools and agent scaffolding have since pushed them much further. Each generation has made the list of things which supposedly require a different kind of intelligence slightly shorter.

The Navier–Stokes system makes the label harder still. What exactly are we classifying? The model weights? The model with a browser and code execution? A coordinated population of ten thousand instances exchanging results over four days? The full research organisation which chooses the prompts and reallocates compute?

Those questions are interesting, but the label matters less than the capability. We do not need consensus on the word AGI to notice that an AI system has apparently produced new research-level mathematics at a scale and speed which would have sounded implausible very recently.

And the model which did it is, according to OpenAI, more capable than Astra and still in training.

What actually ends

I do not think this is the end of mathematics.

There will always be more definitions to invent, more structures to study and more questions to ask. A machine proving theorems does not make those theorems meaningless any more than a telescope discovering a galaxy makes astronomy meaningless. If anything, the amount of mathematics available to us may increase dramatically.

What may be ending is a particular age of mathematics: the period in which the pace of the frontier was constrained by the number of exceptional human minds capable of pushing it forward.

For centuries, the hardest problems acted partly as measures of us. We knew they were difficult because generations of very clever people had failed to solve them. Their eventual solution was expected to tell us something about mathematics and something about the person who had seen what everyone else missed.

If OpenAI’s proof survives scrutiny, Navier–Stokes will still tell us something profound about fluid equations. It may tell us something uncomfortable about ourselves as well.

The frontier is still open. What changed this week is that we can no longer safely assume its pace will be set by human mathematicians.