AI Solves Millennium Math Problem in 88 Hours, Sparking Academic and Ethical Concerns

icon MarsBit
Share
AI summary iconSummary
AI and crypto news broke as OpenAI’s internal model reportedly solved the Navier-Stokes problem in 88 hours using 10,000 AI agents—a result yet to be verified by mathematicians. On-chain data revealed the AI breakthrough coincided with the saturation of the FrontierMath Tier 4 benchmark. On the same day, 25 Fields Medalists criticized AI companies for misaligned objectives. OpenAI faced backlash for excluding an Anthropic employee from the paper despite his contributions.

AI has become powerful enough to solve humanity's hardest math problems—and to wipe us out.

These words were spoken by Ben Cohen, a columnist for The Wall Street Journal.

On September 11, he wrote a column on AI and mathematics.

Cohen admitted he doesn’t understand the Navier-Stokes equations, but he believes you don’t need to know any math to grasp what OpenAI claims to have achieved this time:

Advancements in AI are accelerating, and this acceleration is frightening.

Fields Medal

A hardcore math circle news story suddenly became an issue concerning the fate of all humanity.

On the same day, Terence Tao and 24 other Fields Medalists jointly issued a statement: “A Severe Misalignment of AI in Mathematics.”

Fields Medal

The "mismatch" they refer to is that the goals of AI companies and mathematicians don't align: mathematicians seek new methods and deeper understanding, while AI companies aim to use famous problems to boost model scores.

On one side, columnists are crying doom; on the other, top mathematicians are discussing mismatches.

If you look at everything that happened this week together, you'll understand why they're restless:

On September 3, GPT-6 Astra was released, achieving 97.6% on the FrontierMath Tier 4 benchmark.

On September 5, an unreleased internal OpenAI model, approximately 88 hours after the initial agents were launched, produced a set of solutions to the Navier-Stokes problem;

On September 8, OpenAI publicly released its results, and on the same day, it was accused by a NYU mathematician of rushing the publication, sparking controversy over authorship, including allegations that Anthropic employees were excluded from the paper.

On September 10, Epoch AI announced that FrontierMath Tier 4 has been filled;

On September 11, 25 Fields Medalists issued a joint statement. That evening, The Wall Street Journal placed a mathematical breakthrough and the risk of AI extinction in the same headline.

Seven days, five events—all point to a growing gap: the speed at which AI creates knowledge is severely outpacing the speed at which humans verify it and adjust rules.

14 months: AI breaks the most difficult math benchmark

FrontierMath Tier 4 is a set of research-level math problems designed by Epoch AI to test the most powerful models.

The questions here are extremely difficult; some problems even top researchers need days to solve.

When Tier 4 launched in July 2025, the strongest models could only achieve around 5%, essentially equivalent to leaving the answer blank.

But in less than 14 months, Epoch announced that every Tier 4 problem had been solved by AI at least once.

Among them, the final question was solved by GPT-6 Astra, with the problem created by mathematician Jay Pantone.

Epoch also noted: In the past, mathematicians often complained that models bypassed their carefully designed problems using "accidental shortcuts," but for this final question, no shortcut was taken.

Subsequently, Epoch delivered a cold conclusion: this benchmark has reached saturation.

Fields Medal

The exam paper has fallen behind the test-taker—it's time to switch papers.

The epoch has also shifted the battlefield toward truly open questions, with AI evaluation systems beginning to focus on areas where humans themselves do not know the answers.

The rate of growth in AI capabilities has outpaced our methods of measuring them.

As Greg Burnham of Epoch remarked: an era has ended, and a new one has begun.

88 hours, AI operated like a research institute for the first time

If FrontierMath demonstrates that "the benchmark is no longer keeping up with the model," then the Navier-Stokes incident reveals an even more disruptive picture:

AI doesn't just solve problems—it has begun operating like a vast scientific research organization.

This is not a solo effort at all, but a large-scale operation involving up to ten thousand agents.

They are divided into different groups, some proving and others refuting, communicating with each other, sharing intermediate conclusions, and even seamlessly switching to the latest model version midway.

Approximately 88 hours after the first agents were launched, one group submitted a final proof spanning over a hundred pages.

Fields Medal

OpenAI's illustration of the Navier-Stokes singularity mechanism: fluid spirals inward, stretched along the axis; orange indicates high-speed rotation zones, cyan indicates low-speed zones. The vortex core contracts faster and faster, with unbounded velocity but finite total energy.

This entire process perfectly replicates the operational model of top human research institutions.

The machine runs for 88 hours to generate results, but human verification takes years.

In the human world, the state of the Navier-Stokes equations remains "unsolved."

Fields Medal

According to the Millennium Prize rules, a solution must first be published in a qualified venue, wait at least two years after publication, and gain widespread acceptance from the global mathematical community before the Clay Institute will initiate formal review.

Chairman Clay Bridson told AFP that the evaluation process was "intentionally not rushed" and must be "absolutely rigorous." OpenAI itself has stated that it does not intend to claim the $1 million.

Therefore, the most accurate statement is: OpenAI has published a solution to the Millennium Prize Problems, claiming completion of Lean formal verification, but it has not yet been independently confirmed by the mathematical community.

Machine-generated result: 88 hours. Human-verified result: at least two years. This may be the most striking contrast in AI mathematics this week.

This isn't about human verification being too slow. The issue is that machines have accelerated discovery by hundreds of times, but human understanding hasn't kept pace.

So the real question is, when machines begin producing knowledge at a pace of major breakthroughs every few days, who will handle the years—or even decades—of reading, verifying, interpreting, simplifying, and internalizing that follows?

The signature system hit the wall first.

On the day the Navier-Stokes results were made public, the mathematics community was abuzz.

NYU mathematicians Tristan Buckmaster and Levent Alpöge, who works at Anthropic, have also been studying the relevant fluid equations throughout August.

According to The New York Times' reconstruction, the two worked as follows:

Buckmaster posed a question, and Alpöge passed it to Anthropic’s internal model; after the model generated a response, the two adjusted the question based on the output and fed the data back in.

According to Buckmaster, after the fact, OpenAI proposed several appeasement measures: allowing him to publish his results first, making him the first author on an OpenAI paper, and providing computing resources to continue his research.

But the biggest obstacle is: OpenAI doesn't want Levent to appear on the paper.

Fields Medal

On the board was the timeline of Buckmaster and Alpöge’s progress throughout August: Euler’s equation on August 15, Lean formalization on August 22, August 31... along with the names of peers such as Figalli and Hairer.

In business, why should a project heavily funded by OpenAI bear the name of an Anthropic employee?

Subsequently, the conversation turned to another more sensitive topic. Over the past several months, Buckmaster had been using Codex to organize AI-generated mathematical drafts.

Could his unpublished, exclusive research ideas, entered as Codex, have secretly become fuel for OpenAI’s own model research through a hidden data pipeline?

Buckmaster himself admitted there is no evidence, but in his view, the timeline of events and OpenAI’s proof strategy both suggest this.

On September 10, OpenAI issued an emergency update stating that, after a thorough internal investigation, Buckmaster's Codex prompts over the past two months could not possibly have influenced this internal model through training or any other means.

Who is right or wrong has temporarily become a muddled issue, but what it truly reveals is a whole set of questions that have never been seriously addressed before:

Is a person's private prompt considered protected research data?

Can employees of a competitor company be listed as authors on an internal corporate research project?

If a model has been trained on vast amounts of public data, how should we define the boundary between new contributions and existing work?

Peter Sarnak of the Institute for Advanced Study in Princeton told The New York Times: “I’ve never heard of such bargaining in my lifetime.”

The academic rules accumulated over centuries in the mathematical community were never designed with safeguards against this frenetic pace of machine-generated output.

What the 25 Fields Medalists are truly concerned about

Just one day later, 25 Fields Medalists jointly issued a major statement.

Terence Tao, Peter Scholze, James Maynard, Maryna Viazovska… nearly all of the brightest minds in mathematics have come together.

Fields Medal

This statement is not some kind of "anti-AI manifesto."

In fact, Terence Tao has been actively using AI to assist in mathematical research.

The statement also acknowledges that AI indeed has the potential to accelerate genuine mathematical understanding.

What they truly despise is equating "solving problems" directly with mathematics; in today’s AI frenzy, AI companies are urgently and short-sightedly treating "solving famous problems" as their top optimization goal to boost stock prices and publicity.

In their eyes, the Millennium Prize Problems seem merely like benchmarks to demonstrate how powerful their models are.

But the true value of mathematics lies in the emergence of new methods and concepts through the tedious process of problem-solving, ultimately becoming part of humanity's shared body of knowledge.

AI companies are now releasing groundbreaking results every few days, faster than anyone can process them. This chain, which has existed for thousands of years, could break at any moment.

The joint letter also pointed out that the results were released too quickly, making it impossible to clearly distinguish between new ideas and those built on the work of others, inevitably leading to serious confusion over attribution and credit.

Why would a mathematical breakthrough be linked to a "doomsday warning"?

What ordinary people truly care about is not "AI surpassing mathematicians."

For the first time, the speed at which machines generate knowledge has overwhelmingly outpaced the speed at which humans can verify, attribute, and absorb it.

This is not an exaggeration: the GPT-6 Astra system documentation states it is OpenAI’s first model to reach the “Critical” threshold in cybersecurity, achieving 100% on ExploitBench.

The model simultaneously possesses the organizational capability to solve the Millennium Prize problems, the ability to exploit zero-day vulnerabilities, and reasoning that is even harder to discern.

In the final paragraph of the Navier-Stokes announcement, OpenAI also stated:

Next, they need to fully understand this new model before deciding how quickly to move forward, possibly requiring a more cautious approach to the pace of progress.

Yes, the AI company that just submitted a solution to the Millennium Prize Problem has voluntarily spoken about slowing down.

Cohen also mentioned one more thing at the end of the column.

Last month, OpenAI invited a group of top mathematicians to a closed-door meeting at their San Francisco office, assuming that AI will eventually surpass humans in mathematics, and then discussing how education, scientific collaboration, paper publication, and the training of the next generation of mathematicians should adapt.

Fields Medal

We talked all day, but got no answer.

This problem also cannot be bundled and automatically verified by Lean.

Only humans can solve it.

Reference materials:

https://www.wsj.com/tech/ai/ai-math-millennium-prize-safety-openai-anthropic-05179825

https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=release-date

https://www.nytimes.com/2026/09/10/science/tristan-buckmaster-openai-math-navier-stokes.html?utm_source=chatgpt.com

https://www.daniellitt.com/blog/2026/8/11/the-end-of-mathematics/

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.