Did OpenAI Solve Navier–Stokes? The Bigger Story Is How Humans May Be Teaching AI to Solve Problems

OpenAI's apparent Navier–Stokes breakthrough involved roughly 10,000 AI agents and 2.7 million messages. A dispute over researchers' private AI interactions raises a larger question: when experts work through difficult problems with AI, are they creating something more valuable than answers—a machine-readable record of how experts think?
Person working at a desk facing large screens showing fluid dynamics and AI research workflow diagrams, with equations, proofs, and research notes around them.
Contents

OpenAI has produced a proof that appears to resolve the Navier–Stokes Millennium Prize problem, but the viral version of the story leaves out the most consequential part.

The breakthrough was not simply an AI being handed an impossible math problem and returning an answer 88 hours later. OpenAI describes an organized research operation involving roughly 10,000 concurrent AI agents, easier precursor problems, competing approaches, intermediate discoveries passed between groups, human decisions about where to direct resources, and eventually a formally verified proof. The Clay Mathematics Institute has since said the problem has “apparently been settled,” although its formal prize process will take considerably longer. OpenAI’s account of the Navier–Stokes project (OpenAI)

At the same time, NYU mathematician Tristan Buckmaster raised an uncomfortable question. He and mathematician Levent Alpöge had spent months working on closely related fluid-dynamics problems with extensive assistance from AI systems, including OpenAI’s Codex. Buckmaster says they had placed drafts of their research into those systems. After OpenAI suddenly arrived at the neighboring Navier–Stokes result, he asked whether their earlier interactions might have been used in training. Buckmaster’s full public statement

There is now an important answer to that specific allegation: OpenAI says no.

After investigating, OpenAI said Buckmaster’s Codex prompts from the preceding two months could not have influenced the system in any way, including through training. Buckmaster himself had carefully stopped short of alleging otherwise, writing that he did not know whether his data had been used. There is currently no public evidence establishing that it was. (OpenAI)

But that does not make the controversy unimportant.

It exposes a much larger question that extends far beyond mathematics:

When experts use AI to discover something new, are they merely receiving help from the machine—or are their corrections, failed attempts, judgments and research decisions also creating a record that could teach an AI how experts solve problems?

The distinction matters because an expert’s final answer may not be the most valuable information they produce.

The path to the answer may be worth considerably more.

What exactly did OpenAI solve?

Navier–Stokes equations describe the motion of fluids. They are foundational to areas ranging from aerodynamics and weather modeling to blood flow and engineering.

Instead of following every individual molecule, the equations treat a fluid as a continuous medium and describe properties such as velocity and pressure mathematically.

That works extraordinarily well.

The unresolved problem was whether, in three dimensions, initially well-behaved Navier–Stokes solutions must always remain mathematically smooth—or whether the equations can reach a singularity, a point where quantities such as velocity become unbounded in finite time.

The Clay Mathematics Institute turned that question into one of its seven Millennium Prize Problems in 2000, with $1 million allocated to a valid solution. Clay’s official Navier–Stokes problem description

The official formulation is important because it allows multiple ways to settle the problem.

Two possibilities, known as A and B, would establish global smoothness for the unforced equations.

Two others, C and D, allow a solver to demonstrate breakdown under a suitably smooth external force.

OpenAI says its construction proves those latter cases.

Its system starts with a smooth fluid at rest and applies a smooth external force. A vortex then becomes increasingly concentrated until the velocity becomes unbounded in finite time, while the system retains finite energy. (OpenAI)

That technical distinction matters.

This does not mean airplanes or weather models suddenly stopped working

A singularity permitted by the mathematical equations under a carefully constructed forcing condition does not mean every practical Navier–Stokes simulation is unreliable.

It means something much narrower and mathematically profound:

The official equations do not guarantee globally smooth solutions under every condition allowed by the Millennium Problem formulation.

The counterexample can therefore settle the mathematical question without implying that everyday engineering applications are fundamentally broken.

That nuance is largely disappearing in viral summaries of the story.

Has OpenAI actually won the $1 million?

No.

The Clay Mathematics Institute went unusually far on September 11, saying it was contemplating an announcement that the problem had “apparently been settled.” Clay Mathematics Institute’s September 2026 announcement (Clay Mathematics Institute)

That is substantial recognition.

It is not yet a prize award.

Under Clay’s formal Millennium Prize rules, a proposed solution normally must appear in a qualifying publication, at least two years must pass, the result must achieve general acceptance within the global mathematics community, and Clay must determine that it satisfactorily answers the official problem. (Clay Mathematics Institute)

OpenAI also says it does not intend to claim the Millennium Prize itself. (OpenAI)

So the accurate formulation as of September 2026 is:

OpenAI has published a proof and formal verification that Clay says apparently settles Navier–Stokes, but the formal mathematical and prize-review process is not finished.

The “88 hours” story is more complicated than it sounds

OpenAI says the effort began September 1 after researchers heard rumors that major progress had been made on Millennium Prize problems.

But the AI was not simply given one giant prompt.

OpenAI created groups of agents that pursued different versions of the problem.

It also assigned agents supposedly easier related problems.

One of those involved the Euler equations—the Navier–Stokes equations with the viscosity term removed.

Nearly 100 agents worked for about 50 hours on that problem and produced an unforced Euler blowup result. (OpenAI)

Humans then made a crucial decision.

Because the Euler result looked promising, OpenAI shifted resources away from other Millennium Problems and toward Navier–Stokes. The new agent groups were given the Euler result as additional information.

Different groups explored different approaches.

OpenAI says it then used Codex to consolidate their most useful discoveries and cross-pollinate the groups.

The Navier–Stokes result arrived roughly 88 hours after the first agents were launched. Formalization and verification in the Lean theorem prover took another approximately 17 hours. (OpenAI)

The scale is difficult to overstate.

OpenAI reports that the Navier–Stokes portion alone involved approximately:

10,000 concurrent agents, 2.7 million agent messages and 130 billion output tokens. (OpenAI)

That is not merely a more intelligent chatbot.

It resembles a computational research organization.

And that may be the most important technical development in the entire story.

Two mathematicians were already following a closely related path

The human side begins earlier.

Buckmaster and Alpöge had been building on work by mathematicians Diego Córdoba and Luis Martínez-Zoroa, who had developed a program for constructing fluid blowups using external forcing.

Buckmaster is explicit about the intellectual lineage. Their program was not invented by an LLM; he credits Córdoba and Martínez-Zoroa with the foundational approach.

Buckmaster and Alpöge then used multiple AI systems extensively while attempting to push that strategy further.

According to Buckmaster, they used Claude and OpenAI’s Codex throughout their work. By August 15, they had obtained smooth-forcing blowup results for Boussinesq and three-dimensional incompressible Euler, and they verified the Euler result in Lean on August 22.

NYU subsequently described their work as a significant advance toward Navier–Stokes and confirmed that large language models assisted in producing and verifying the proofs. NYU Courant’s description of Buckmaster’s work (NYU CIMS)

Then rumors about their progress began circulating.

OpenAI began its own effort.

Within days, its system produced the stronger Navier–Stokes result.

That timing created the controversy.

Did Buckmaster and Alpöge accidentally teach OpenAI how to solve it?

There is currently no evidence that they did.

That distinction needs to be explicit because some online versions of the story have already crossed from a legitimate question into an unsupported conclusion.

Buckmaster said that during the project he and Alpöge had been putting drafts into Codex.

When he learned that OpenAI’s system had reached a smooth-forcing Navier–Stokes result, he asked whether the internal model had been trained on or otherwise had access to those Codex sessions.

He wrote that he initially received an answer about the model not looking up user data but did not receive an answer to his separate question about training.

Buckmaster then made his own uncertainty unusually clear:

He had not seen OpenAI’s proof at the time. He did not know how its system obtained the result. He did not know whether their data had been used, and he said he was not accusing anyone of doing so.

OpenAI later published a stronger answer.

Following an investigation, the company said Buckmaster’s Codex prompts during the two months before the announcement could not have influenced the system, including through training. OpenAI said the system was based on a previously pretrained model followed by large-scale reinforcement learning and said its proof also differed substantially from the researchers’ Euler result. (OpenAI)

Absent contradictory evidence, that is the best-supported conclusion.

But answering that factual question leaves the more interesting one untouched.

Can working with an AI teach the AI how you think?

Potentially—but the word teach hides several completely different mechanisms.

That distinction is essential.

1. You can teach an AI inside the current conversation

Suppose an AI initially analyzes something badly.

You tell it:

That source is repeating a press release. Find the underlying document.

It searches again.

Then you tell it:

You’re treating two pieces of evidence as equivalent even though one is a primary record.

It changes its analysis.

Then:

The timeline contradicts your explanation. Reconstruct it chronologically.

Again, the output improves.

The underlying neural-network weights do not need to change for any of this to happen.

The model can learn from information supplied inside its current context.

This phenomenon is known as in-context learning.

Google DeepMind researchers, for example, have demonstrated that models can learn from hundreds or thousands of examples presented within their context and substantially improve at new tasks without updating the model’s weights. Google DeepMind’s research on many-shot in-context learning (Google DeepMind)

That is temporary learning in an important sense.

The model has been taught how to handle the immediate task, but the underlying general model has not necessarily been retrained.

2. A system can retain methods through memory, instructions or retrieval

There is another layer between a single conversation and foundational model training.

An AI system can preserve:

  • project instructions;
  • editorial rules;
  • examples;
  • documents;
  • previous decisions;
  • databases;
  • retrieval indexes;
  • reusable workflows.

Those mechanisms can make the system behave as though it has permanently learned a methodology even when the foundation model’s parameters remain unchanged.

For a user, that distinction can be almost invisible.

The AI seems to have learned.

Technically, however, it may simply be retrieving the instructions you previously gave it.

3. Your interactions can potentially become future training material

This is the mechanism at the center of the Buckmaster concern.

OpenAI’s current data-use policy states that content from individual consumer products may be used to improve and train models unless the user opts out.

The same policy says that ChatGPT Business/Team, Enterprise and API inputs and outputs are not used for model training by default unless an organization explicitly opts in. (OpenAI)

That does not mean every eligible conversation becomes training data.

It also does not mean a model learns something immediately after someone types it.

Training datasets have to be collected, processed and incorporated into later training or post-training runs.

But the mechanism exists.

An eligible interaction today can, depending on the provider’s policies and training process, potentially contribute to a future model.

4. Publishing the result creates another pathway entirely

Once research becomes public, models may be able to obtain it through web browsing, retrieval systems or future training corpora.

That is different from using private conversations.

An AI retrieving a publicly posted mathematical paper tomorrow does not prove that it was trained on the paper today.

Those mechanisms are routinely confused in discussions about AI.

They should not be.

The final answer may be less valuable than the path to it

Here is where this story becomes much bigger than Navier–Stokes.

Imagine two datasets.

The first contains a finished mathematical proof.

The second contains the months of work that led to it:

  • promising ideas;
  • failed approaches;
  • corrections;
  • objections;
  • mistaken derivations;
  • literature searches;
  • discarded hypotheses;
  • expert judgments;
  • revised proofs;
  • explanations of why one route failed and another worked.

Which dataset better teaches another researcher how to solve similar problems?

Very possibly the second.

AI research already gives us a reason to take that distinction seriously.

In a well-known study of mathematical reasoning, OpenAI researchers compared outcome supervision—rewarding a model for reaching the correct final answer—with process supervision, which evaluates intermediate reasoning steps.

The process-supervised approach performed substantially better in their experiment. OpenAI’s research on process supervision (OpenAI)

That does not prove that researchers’ chat histories are being converted into process-supervision datasets.

It demonstrates something more fundamental:

Information about how a problem is solved can itself be highly valuable training data.

A sherafy.com investigation makes the difference obvious

Consider a typical sherafy.com research process.

Someone posts a viral claim.

The finished article might eventually contain:

Claim → evidence → explanation → conclusion.

That final article is useful.

But imagine giving an AI the entire research history instead.

The first response misunderstands the issue.

The editor rejects it.

The AI searches again.

A secondary report is discovered.

The editor says to find the primary filing.

The filing contradicts part of the secondary coverage.

The AI proposes three explanations.

Two are rejected because the chronology does not fit.

Another source appears to confirm the remaining explanation, but the editor notices that both stories trace back to the same original source, so they are not independent corroboration.

The research gets redirected.

A numerical claim is recalculated.

A distinction between allegation and verified fact is added.

Then the evidence is recursively reconsidered until the surviving explanation accounts for the largest number of facts with the fewest unsupported assumptions.

Finally, an article is written.

The published article contains the result of that work.

The conversation contains something else:

the editorial algorithm.

Not software code.

A reasoning procedure.

It records things such as:

what deserves investigation → which sources deserve greater weight → what counts as corroboration → how contradictions are handled → when an explanation should be rejected → when additional research is necessary → how uncertainty is expressed → when inference is justified.

That is far richer than the final prose.

One correction can contain more useful information than ten accepted answers

Suppose the AI proposes:

Two newspapers reported the number, so the claim is corroborated.

The editor replies:

No. Both newspapers got the figure from the same report. That’s one source repeated twice. Find independent evidence.

That short correction teaches a generalizable rule about evidence.

Or:

You’re repeating what the agency says happened. Check whether the underlying court filing actually supports it.

Again, that is more than an answer.

It is supervision about how research should be performed.

Or:

Your explanation works until August 12. The August 9 document means the decision had already been made. Rebuild the chronology.

Another reusable reasoning pattern.

Do that thousands of times and you have something resembling a corpus of expert judgments about investigation.

This does not mean a model automatically absorbs them.

But if those interactions were deliberately converted into training examples, they could contain precisely the sort of expert feedback used to improve an AI’s future behavior.

This may be the hidden value of widespread AI adoption

Most discussions about AI training focus on content.

Books.

Websites.

Code.

Photos.

Scientific papers.

News articles.

Those contain enormous amounts of human knowledge.

But conversational AI is generating another category of data that may be even more interesting:

humans correcting machines while doing real work.

A lawyer tells the AI why its interpretation of a contract is wrong.

A physician identifies the symptom that actually matters.

An engineer rejects an unstable design.

A programmer diagnoses why the proposed patch fails.

A mathematician explains why a proof strategy cannot work.

An investigative editor tells the system why three apparent sources are really one.

Every correction potentially reveals information about expert judgment that the finished product may never contain.

That does not mean providers necessarily use all such interactions, nor that every product permits them to do so.

It means the data itself can be unusually valuable.

The Navier–Stokes result also reveals a second revolution

There is another part of OpenAI’s account that may ultimately matter even more than the dispute over training data.

For years, the central AI question has been:

How intelligent can one model become?

The Navier–Stokes experiment suggests another question:

What happens when thousands of copies of that intelligence can be organized into a research institution?

OpenAI did not merely scale inference by asking the same question 10,000 times.

According to its account, the system divided labor.

Different groups attacked different formulations.

Some worked on easier precursor problems.

Intermediate discoveries were evaluated.

Resources were redirected.

Successful ideas were shared among groups.

The system recursively narrowed the search space until one route worked. (OpenAI)

That looks less like a chatbot.

It looks more like a research organization whose workers can be duplicated nearly without limit.

The human role did not disappear

In fact, OpenAI’s account demonstrates the opposite.

Humans decided:

  • which problems to include;
  • how to divide them;
  • which easier problems might unlock the harder one;
  • when the Euler result warranted shifting resources;
  • what intermediate findings should be propagated;
  • how to organize the agents;
  • when formal verification was required.

The AI performed enormous amounts of reasoning.

Humans designed the search process through which that reasoning became useful.

That distinction points toward a different division of intellectual labor.

The valuable human may increasingly be the person who knows:

what to investigate, how to decompose it, what evidence matters, which result is suspicious, where the machine is likely wrong, and when to redirect thousands of machine-hours toward a better hypothesis.

That applies to mathematics.

It also applies to journalism.

What would 10,000 research agents look like at sherafy.com?

Take the same architecture and apply it to an investigation.

Instead of assigning one AI to “research this story,” divide the problem.

One group reconstructs the timeline.

Another traces ownership.

Another finds primary documents.

Another audits numerical claims.

Another searches historical precedents.

Another looks specifically for evidence contradicting the leading hypothesis.

Another traces every secondary claim backward to its original source.

Another identifies missing information.

Another attempts to falsify the draft conclusion.

A separate system compares their findings and detects contradictions.

The editor does not read 10,000 raw outputs.

The agents recursively compress them.

Weak branches die.

Promising branches receive more resources.

Eventually the strongest evidence reaches the human.

The process becomes:

Human question → parallel machine investigation → machine synthesis → human judgment → targeted machine investigation → adversarial verification → final analysis.

That is conceptually close to what OpenAI describes doing with Navier–Stokes.

And it suggests that the next stage of AI-assisted knowledge work may not be humans asking better questions of individual chatbots.

It may be humans designing research organizations made of AI agents.

The strange inversion: using AI can reveal how experts think

There is an irony here.

AI was originally trained by ingesting enormous amounts of the output humans had already produced.

Books taught it prose.

Repositories taught it code.

Papers taught it science.

Websites taught it explanations.

But conversational AI potentially exposes something traditional publications rarely preserve:

the mistakes and corrections before publication.

A published sherafy.com article does not show every bad hypothesis that was rejected.

A finished mathematical paper does not contain every failed proof.

A piece of production software does not preserve every wrong architectural decision.

An AI conversation can.

That makes the interaction record qualitatively different from the finished work.

It can encode how a knowledgeable human distinguishes good reasoning from bad reasoning.

And that may become an increasingly valuable resource for improving future AI systems.

But we should not turn that possibility into another unsupported conspiracy

This is where the Navier–Stokes story requires unusually careful language.

Several things are established.

Buckmaster and Alpöge used AI extensively.

They uploaded drafts into Codex.

OpenAI later launched an intensive effort after hearing rumors about their work.

The two efforts pursued related fluid-blowup research.

Buckmaster asked whether his Codex interactions had affected OpenAI’s model.

OpenAI subsequently investigated and says the relevant prompts could not have influenced the model, including through training.

Those facts justify asking broad questions about AI research tools and data governance.

They do not justify stating that OpenAI secretly trained on Buckmaster’s unpublished proof.

In fact, the broader argument becomes stronger once that unsupported leap is removed.

We do not need this particular controversy to prove improper data use.

The established facts already demonstrate the underlying transformation.

Researchers are increasingly thinking through AI systems, rather than merely typing finished work into them.

Those interactions can contain detailed records of expert reasoning.

AI companies have technical mechanisms for learning from human feedback.

Some consumer AI interactions can, depending on policies and settings, be eligible for future model improvement.

And the Navier–Stokes experiment demonstrates that sufficiently capable models can turn knowledge, intermediate discoveries and massive parallel search into genuinely new research.

Every part of that chain is real.

The biggest lesson may not be that AI solved Navier–Stokes

The immediate historical headline is extraordinary:

A machine system appears to have resolved a mathematical problem that resisted generations of mathematicians.

But the more consequential development may be hiding underneath it.

Humans are beginning to externalize their problem-solving processes into AI-readable form.

Every time an expert says:

No, that approach fails because…

This source is stronger because…

Try this formulation instead…

You’re missing the contradiction here…

Go back three steps…

That result changes which problem we should attack next…

they are expressing expertise that previously remained largely inside the person’s head.

The finished work records what humans discovered.

The interaction history can record how they discovered it.

That is a profound difference.

And as systems like OpenAI’s Navier–Stokes experiment move from one AI assistant toward thousands of coordinated agents, human intellectual leverage may increasingly come from deciding what those machines should explore, what they should ignore, which discoveries matter, and when their conclusions should not be trusted.

The central question therefore may not be:

Can AI replace the expert?

A better question is:

What happens when experts spend years converting the parts of expertise that were once tacit into instructions, corrections and research trajectories that machines can understand?

Navier–Stokes may eventually be remembered as an important mathematical breakthrough.

It may also be remembered as an early warning that the most valuable thing humans give AI systems is not always the answer.

Sometimes it is the method for finding the next one.

References and Further Reading

Navier–Stokes and the 2026 result

On the Navier–Stokes Millennium Prize Problem — OpenAI — OpenAI’s September 8, 2026 account of the proof, multi-agent research process, computation used, relationship to the concurrent Euler work and its subsequent investigation of Buckmaster’s Codex interactions. (OpenAI)

Navier–Stokes Announcement — Clay Mathematics Institute — Clay’s September 11 statement saying the Millennium Problem had “apparently been settled” while emphasizing that formal evaluation would proceed under its established rules. (Clay Mathematics Institute)

Existence and Smoothness of the Navier–Stokes Equation — Charles L. Fefferman / Clay Mathematics Institute — The official mathematical formulation of the problem, including the A, B, C and D routes by which it can be resolved.

Millennium Prize Description and Rules — Clay Mathematics Institute — Governs publication, the minimum two-year waiting period, general mathematical acceptance and Clay’s evaluation before a $1 million prize can be awarded. (Clay Mathematics Institute)

Buckmaster, Alpöge and the concurrent research

Tristan Buckmaster’s September 2026 Statement — Buckmaster’s first-person chronology of the Euler work, use of Claude and Codex, interactions with OpenAI, his questions concerning training data and his explicit statement that he did not know whether their data had been used.

Tristan Buckmaster Constructs Finite-Time Blowup for 3D Euler with Smooth Forcing — NYU Courant — NYU’s account of the mathematical advance, its intellectual lineage and the role of large language models and Lean verification. (NYU CIMS)

How AI can learn from humans

Many-Shot In-Context Learning — Google DeepMind — Research demonstrating that large language models can learn substantially from examples placed in their active context without updates to their underlying model weights. (Google DeepMind)

Improving Mathematical Reasoning With Process Supervision — OpenAI — Research comparing feedback on final answers with feedback on intermediate reasoning steps, illustrating why records of a problem-solving process can contain training value beyond the final outcome. (OpenAI)

How Your Data Is Used to Improve Model Performance — OpenAI — OpenAI’s current explanation of when individual consumer content may be used for model improvement and its default no-training policy for business, Enterprise and API inputs and outputs. (OpenAI)

Editorial currency note: The Navier–Stokes proof, Clay’s review status, AI product names and provider data-use policies are evolving. This article reflects publicly available information reviewed through September 16, 2026.

Cite this article

Published September 16, 2026

Think something here is wrong, incomplete, outdated, or insufficiently supported? You can challenge a factual claim, source, interpretation, missing context, or privacy issue.

Learn How the challenge process works


More to think on...

A desk setup with multiple monitors showing AI agent workflow, audit logs, and billing data highlighting $5,427.31 in charges.
Did GPT-6 Astra Really Spend $5,000 on Unauthorized Seedance Videos? What the Evidence Shows

Sirio Berati says GPT-6 Astra turned a 59-video Seedance job into roughly 500 generations, ran up more than $5,000 in charges, and then gave him an inaccurate account of what happened. The underlying allegation is technically plausible, but several viral claims—including that Astra “hacked” his account or deliberately rotated IP addresses to hide its actions—are not yet established by public evidence.

Read More »