Anthropic's Claude proves Fermat's Last Theorem in 11 days using 13 million lines of code.

icon MarsBit
Share
AI summary iconSummary
Anthropic's Claude AI achieved a major breakthrough in AI and crypto news by completing the first fully machine-verified proof of Fermat’s Last Theorem in 11 days. Under the leadership of Peng Tianyi from Tsinghua University, the AI generated 13 million lines of code and verified 29,500 theorems—surpassing the scale of Mathlib, the largest mathematical theorem library. This advancement underscores AI’s expanding role in solving complex problems, providing powerful new tools for researchers and developers in the crypto space.

Just now, there was shocking news from the math circle.

Led by a top student from Tsinghua's Yao Class, Claude has completely solved Fermat's Last Theorem.

Thus, AI has completed the largest proof in the history of mathematics.

Once, Fermat's Last Theorem tormented humanity for over 350 years, requiring mathematicians to spend years of dedicated work to write a 129-page proof.

Today, Anthropic announced that Claude completed the first end-to-end machine-verified proof of Fermat's Last Theorem in just 11 days!

Fermat's Last Theorem

To achieve this, Claude wrote 13 million lines of code, generating 30,300 verifiable theorems, of which 29,500 were adopted directly into the final proof.

This scale is more than five times that of the world’s largest mathematical theorem library, Mathlib—and the entire process consumed a staggering 6 billion tokens.

This is the largest Lean proof ever written.

As soon as the news broke, the entire internet was shaken. Some exclaimed, “Formalized Fermat’s Last Theorem in just one month? This display of strength makes the mathematical community look like snails crawling.”

Fermat's Last Theorem

Peng Tianyi, a standout from the Yao Class

A 350-year-old century-old problem solved by Claude in 11 days

In 1637, the French mathematician Fermat, while reading a book, casually wrote in the margin:

When the integer n > 2, the equation xⁿ + yⁿ = zⁿ has no positive integer solutions for x, y, and z.

Then he added, "I have discovered a truly marvelous proof, but this margin is too narrow to contain it."

This single sentence tormented mathematicians for over three hundred years.

It wasn't until 1995 that British mathematician Wiles used advanced modern mathematical tools to produce a 129-page proof paper, finally resolving this 350-year-old mystery.

Fermat's Last Theorem

But the problem is: Wiles's proof is extremely complex.

Modern mathematics has advanced to the point where ordinary people can't even understand the problems. Proving a theorem is like building an extremely complex chain of logic—break just one link, and the entire structure collapses.

When Wiles first announced his proof in 1993, a critical flaw was discovered, and he spent another year in intense, solitary effort to fix it. For such a high-level mathematical proof, verifying its correctness often requires top experts to spend months or even years.

Is there a way for a computer to verify the correctness of a result simply by running through it once, just like checking the output of a calculator?

Yes! This is "formalization."

In simple terms, it means translating human-written mathematical proofs into a programming language that computers can execute (such as Lean), then letting the machine verify each step logically; if the process completes successfully, the proof is guaranteed to be correct.

However, formalizing Fermat's Last Theorem is widely regarded by the mathematical community as a massive undertaking that takes years.

The first-phase blueprint of the project led solely by Imperial College professor Kevin Buzzard is 86 pages long!

Fermat's Last Theorem

Then, Claude arrived.

What humans were expected to take years to accomplish was completed in just 11 days—and it was done with “basic autonomous operation.”

13 million lines of code, 6 billion tokens

"11 days, 13 million lines of code"—behind this lies the dual impact of AI’s brute-force aesthetics and precision system design.

Let’s see what Claude actually did:

It not only proved Fermat's Last Theorem itself.

Because formal proofs must begin from the most fundamental axioms and build upward layer by layer, Claude also proved the more than 29,000 other mathematical theorems required along the way.

It involved algebra, geometry, number theory, harmonic analysis—many branches that had never been formalized before, and Claude essentially pioneered them.

Throughout the entire process, human involvement was minimal.

Researchers provided some high-level directives, such as "The Jacobian variety as a scheme is a high priority" and "Accelerate progress on Mazur's theorem."

Fermat's Last Theorem

The rest involves dozens of Claude agents wildly conversing with each other, defining concepts, proving intermediate theorems, and building upward layer by layer.

Finally, the Lean compiler passed a full check, relying only on three fundamental axioms.

When the program finished and the console displayed the sacred message "PROVED," Claude's own internal logs were excited:

!!! Fermat's Last Theorem root node read as PROVED... This was the objective of this campaign... A historic moment.

Look, even AI knows how amazing this is.

The mastermind behind the scenes: a super-achiever from Tsinghua University's prestigious Yao Class

Only someone extraordinary could command Claude to perform such a miracle.

Leading the group is Tianyi Peng, Assistant Professor at Columbia Business School and researcher at Anthropic.

His resume is basically a cheat code:

Bachelor's degree, 2013–2017, Tsinghua University's "Yao Class"; recipient of the outstanding thesis award; selected for the National Informatics Olympiad Training Team.

The doctor went to MIT, studied operations research, and graduated with a perfect GPA of 5.0.

Currently serving as an assistant professor at Columbia University while working on AI agents and formal tools at Anthropic.

Interestingly, Peng Tianyi's obsession with "AI automatically verifying mathematical proofs" stems from a painful experience during his undergraduate years.

At the time, his advisor wanted to submit the results from his thesis to Nature, but asked him: “Are you 100% certain the proof is correct?”

He answered honestly: “About 99% confident, but with something this long, I can’t be 100% certain.”

Because of that 1% uncertainty, he missed the opportunity to publish in Nature.

Now it’s fixed—he used AI to shut down that “1%” himself.

From Flop to Fame: How Prove2Me Saved AI

Do you think feeding AI a single prompt like “Prove Fermat’s Last Theorem” will instantly spit out 13 million lines of code?

Completely wrong.

At first, the experiment nearly failed.

Anthropic revealed that early attempts to have dozens of Claude agents collaborate quickly descended into chaos, like flies without heads, unable to keep up with each other and achieving abysmally low collaboration efficiency. The code contributed from these early failed attempts ultimately accounted for only 7%.

The "forgetfulness" and "hallucinations" of large models are fatal flaws when faced with rigorous mathematics—one error renders millions of subsequent lines invalid.

At a critical moment, Peng Tianyi's team launched the Prove2Me platform.

This is like giving AIs a "super project manager" designed to handle anything.

Theorem DAG (task tree): Provides each AI with a clear map indicating which intermediate node to prove next, significantly alleviating memory decay and enabling dozens of agents to operate efficiently in parallel.

Separation of declaration and proof: Faster compilation and reduced resource usage.

Natural language indexing: Each theorem includes a plain-language description, making it much easier for AI systems to retrieve and reuse the results.

Fermat's Last Theorem

With Claude Code’s multi-agent framework, AIs race through mathematical mazes like engineers with GPS, clearing all levels in 11 days.

The math community has given up.

As soon as the results were released, X and major tech forums were hit by a tsunami.

The professor from Imperial College London, who originally planned to spend years working on formalization, was thoroughly impressed and gave high praise: “This is a major step forward in the automated formalization of modern mathematical literature. It can be used to detect errors in human mathematical libraries and verify mathematical conclusions generated by large models.”

But netizens' imaginations are even more creative.

Someone commented: “How AI solves math problems—by giving you an incredibly complex answer (13 million lines of code). Proving it wrong is harder than solving it yourself, so you’re left with no choice but to give up and accept that it’s right. Isn’t that just PUA in the world of mathematics?”

Others say: “Writing 13 million lines of code just to make a 350-year-old theorem sit quietly in front of a machine… humanity’s table isn’t ready to handle something of this scale.”

Even more impressively, Anthropic conducted a small experiment to showcase its capabilities.

Using three regular accounts, they formalized the famous "Vinogradov's Three Primes Theorem" from number theory on Prove2Me in just three days!

In other words, with the right tools, amateur scientists will soon be able to verify the most advanced mathematical theorems using just a few consumer-grade AI accounts!

AI will not replace mathematicians, but it will radically transform mathematics.

So, are mathematicians going to lose their jobs?

Anthropic officially answered: It won't replace, but it will completely change the game.

Historically, there have been many "tragedies" in mathematical verification.

For example, in 1998, someone proved the Kepler conjecture, and the review panel spent four years before concluding they were “99% certain”; Perelman proved the Poincaré conjecture, and the entire mathematical community spent four years writing three books, each over 300 pages, just to barely understand it; some theorems were accepted as truth for years, with others building entire structures on them, only to later discover the foundation was flawed.

And the technology Claude brings is designed to eliminate this "uncertainty."

In the future, AI will be not just a calculator for mathematicians, but also the strictest judge.

When AI can rapidly generate thousands of new conjectures and proofs that humans cannot possibly review, make "accompanying formal verification code" a standard requirement for papers.

In the era of large models, large-scale automated formalization opens up new possibilities through near-engineering implementation.

Reference materials:

https://www.anthropic.com/research/formalizing-fermats-last-theorem

Edited by: Aeneas

This article is from the WeChat public account "New Intelligence Yuan" (ID: AI_era), authored by ASI Revelation.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.