AI Solves Erdős Problem in One Page, Outperforming a 44-Page Proof

icon MarsBit
Share
AI summary iconSummary
AI solves the Erdős problem in one page, outperforming a 44-page proof from 1991. The problem, listed as number 119 in the Erdős Problems, was verified as correct by Thomas Bloom. GPT-5.6 and mathematician Korsky used elementary harmonic analysis to resolve the result. Altcoins to watch may gain momentum as the Fear & Greed Index shifts in response to AI-driven breakthroughs.

One page conquered a tough problem that would have required 44 pages of a top-tier paper.

In 1991, a 44-page paper was published in the Annals of Mathematics.

The author is József Beck, a name well known in the field of combinatorics.

He is the originator of the Beck-Fiala theorem, a 1985 Fulkerson Prize winner, an invited speaker at the 1986 International Congress of Mathematicians, and one of the founders of discrepancy theory.

The paper title is "The Modulus of Polynomials with Zeros on the Unit Circle: A Problem of Erdős," addressing problem number 119 in the Erdős Problem Collection.

Take any sequence of complex numbers on the unit circle as zeros, multiply them to form a polynomial; can the maximum modulus Mn of this polynomial on the unit circle be guaranteed to exceed n^c for infinitely many n?

Beck proved that it could be done. This problem was widely regarded as the toughest challenge of that era.

Annals of Mathematics

AI claimed Erdős's bounty. Problem #119 originally posed three questions in one, with the third carrying a $100 reward—its status is now marked as SOLVED.

Thirty-five years later, this paper was surpassed by a single page.

These words were spoken by Thomas Bloom.

He is a mathematician, a researcher at the University of Manchester, and the creator and maintainer of erdosproblems.com.

More than a thousand open problems posed by Erdős are listed on this website, serving as the first stop for mathematicians worldwide to track these challenges.

That's why every AI company claiming to have "solved the Erdős problem" must first pass his scrutiny.

He wrote on X:

So far, GPT-5.6 Sol is the most mathematically intriguing new model I’ve encountered. Every new proof submitted to erdosproblems.com that I’ve examined in detail has been correct and contains interesting insights.

Annals of Mathematics

Bloom also emphasized that this is not because AI deployed any more sophisticated tools, but because we originally thought the problem was much harder than it actually was.

OpenAI President Greg Brockman retweeted this post, saying:

This feels like a watershed moment for mathematics. Scientific and medical breakthroughs that can truly improve human life now seem within reach.

Annals of Mathematics

Erdős's $100 prize has been claimed by AI

First, let's clarify the background and details of this question.

In 1957, Erdős posed this problem in a paper, asking three consecutive questions about the modulus of zero polynomials on the unit circle.

Over the next forty years, he revisited it repeatedly in 1961, 1964, 1982, 1990, and 1997.

The first question was won by Wagner in 1980.

The second question: Beck proved in 1991 that there exists some c > 0 such that max_{n ≤ N} M_n > N^c. This was presented in the 44-page paper published in the Annals of Mathematics.

Annals of Mathematics

Beck's paper was published in Volume 134, Issue 3 of the Annals of Mathematics in 1991.

For the third question, Erdős personally offered a $100 reward, as stated in his 1997 paper.

Now, the bounty has been claimed by GPT-5.6 and a mathematician named Korsky, who solved the problem Beck left untouched using just one page.

Annals of Mathematics

Erdős posed over a thousand problems throughout his life and offered cash prizes ranging from $25 to $10,000 for their solutions.

Bloom's evaluation is: Under Korsky's guidance, GPT proved a result stronger than Beck's 44-page Annals paper using just one page of elementary harmonic analysis techniques.

A crucial step, which Bloom later clarified in the discussion forum: convolving a non-negative function and then shifting it in time by one unit immediately smooths the expression, making the previously difficult parts flow naturally.

He said that this approach of estimating the operator norm using the dual norm and inner product is one of the most common techniques in analysis.

Not a new tool, but an old tool used in a place no one had thought of.

Annals of Mathematics

Bloom noted that Beck's paper in the Annals contains many important and interesting ideas, though they are not relevant to this problem. He added, "At least to me, that was surprising."

He also mentioned a detail: no one has ever calculated that constant, and it’s likely very small. Beck himself never calculated it, and according to his judgment, if that constant were significant, Beck would have calculated it.

So, this isn't AI overturning top journals—it's discovering a shortcut humans could have taken all along.

For mathematicians, this is far more exciting than "AI solved another problem."

Solving a problem simply added a tool. This time, the tool corrected the person using it.

Humans might shrug and walk away; AI will try the next one.

If this were the only thing, it would be at best a beautiful anomaly.

GPT-5.6 Sol Ultra was fully released on July 9, and on July 10, OpenAI announced that GPT-5.6 Sol Ultra produced a complete proof of the Cycle Double Cover conjecture.

This problem was independently posed by Szekeres in 1973 and Seymour in 1979, and remained unsolved for about fifty years.

The model was allocated 8 hours but completed in less than 1 hour, using 64 sub-agents running in parallel. The prompts and proof PDFs are fully public.

Bloom was the first to provide a substantive evaluation.

He said it was a beautiful proof and provided three adjectives:

Short. Elementary. It could have been discovered in the 1980s.

He also pointed out the underlying reason: he speculated that there was a small, counterintuitive turn in that crucial step.

How would a human mathematician approach this problem?

He would first try the most natural approach—review the linear algebra, realize it doesn’t work, shrug, think to himself, “I didn’t expect it to be this easy,” and walk away.

But AI won't get discouraged—it will keep trying small variations until one works.

AI excels at avoiding emotional stop-losses.

Looking back at April this year, GPT-5.4 Pro solved Erdős 1196 in 80 minutes using the von Mangoldt function.

Jared Lichtman of Oxford spent seven years on that problem, and later said that everyone working on the original set problem since 1935 had become accustomed to the same opening move—one that obscured a technical possibility that had been right in plain sight for 90 years.

Terence Tao’s evaluation was even harsher: Over the past few decades, humanity as a whole has gone off track right from the first step.

Human intuition is certainly not a flaw. On the contrary, it is an efficiency tool: it saves us from countless doomed attempts, and without it, no one could reach the frontiers of their field within a limited lifetime.

The cost is that occasionally, it also skips the one it shouldn't.

People who said GPT-5.6 is interesting said another thing last year.

In October last year, Bloom publicly contradicted OpenAI.

At the time, OpenAI VP Kevin Weil shared Mark Sellke’s post and wrote: “GPT-5 has just found solutions to 10 previously unsolved Erdős problems, made progress on 11 others, all of which have been open for decades.”

Annals of Mathematics

Bloom's one-line response characterizes this as dramatic misinformation.

He said that marking these problems as "open" on the website only means he personally isn't aware of any papers that have solved them. What GPT-5 did was uncover those papers he wasn't aware of.

Finding literature is not the same as producing proof.

Annals of Mathematics

The following scenario is remembered by many.

LeCun chimed in, sarcastically noting that the bomb he had planted had blown up in his own face. Hassabis was more direct: “This is too embarrassing.”

Annals of Mathematics

Wei deleted the post.

OpenAI researcher Sébastien Bubeck finally admitted that he had merely found a solution already present in the literature.

In May this year, when OpenAI announced the refutation of a 1946 Erdős geometric conjecture, the announcement included evaluations from Noga Alon, Melanie Wood, and Thomas Bloom.

Has AI hit a mathematical wall?

The debate under Bloom's post was also fascinating.

The bearish camp, represented by user qrdl, stated: “The thought chain output for 5.6 is minimal, there’s been no progress on problem 677, and the CritPT evaluation on the physics side has barely changed compared to 5.4.”

His conclusion is: AI has hit a wall in mathematics.

Annals of Mathematics

User qrdl posted on the forum, saying that CritPT has barely moved since 5.4, and unless they find a way to boost the benchmark score, the results will remain the same.

The argument for qrdl is not emotional.

He believes that large models are essentially static-weight token generators that, when doing math, do not form new neural connections in real time like humans do. True creative leaps require a flexible brain.

The experience with qrdl is: every time a new session is started to attack 677, the model repeats the same old tactics that have already failed.

The comeback came just as quickly.

Nat Sothanaphan pointed out that CritPT is a physical test, not a mathematical benchmark; using its single curve to claim that AI has hit a wall in mathematics is cherry-picking favorable data.

His counter-evidence is that 5.6 shows improvement over 5.5 across a wide range of benchmarks, including FrontierMath.

The other camp is the practical camp, represented by the user named old-bielefelder.

Annals of Mathematics

The old-bielefelder report states that GPT-5.6 Sol spent 14 minutes thinking and over 9 minutes reviewing, improving the index from 0.72 in Pintz (2018) to 0.7195.

Not earth-shattering, but it did move—using old-bielefelder’s own words, it’s the “first milk.”

In contrast, the previous night, GPT-5.5 struggled for over an hour on the same question, with 0.72 remaining unchanged.

He then notified Pintz directly of the results.

A third voice, also from Nat Sothanaphan, borrowed Bloom’s phrasing: modern mathematics is a vast cathedral.

Simply reaching the forefront of this church is extremely exhausting, yet large models can now accomplish this with superhuman efficiency—most recent breakthroughs stem from this. Moving further forward requires fluid intelligence.

Some people say large models lack fluid intelligence, and he believes they are wrong.

The sustained price increase of the ARC series over the past few years is evidence of this; GPT-5.6 Sol achieved 7.8% at its highest inference tier. When this benchmark launched in March this year, the best result was 0.37%, while humans consistently maintained above 90%.

Annals of Mathematics

ARC Prize official verification results. GPT-5.6 Sol's performance on ARC-AGI-3 sharply increases with reasoning level: 0.3% at Low setting, 7.8% at Max setting.

His explanation was: the foundation of the church was so massive that it completely overshadowed the small amount of rapidly growing fluid intelligence. So, if you only look at the output, it seems like nothing is happening.

But one day, when fluid intelligence begins to outweigh the foundation, it will suddenly become visible to the naked eye.

After two weeks of debate, what this argument has truly revealed is not whether AI can do math, but how to define the word "difficult."

In the history of mathematics, some entries marked as "unsolved" reflect the inherent difficulty of the problems themselves, while others merely reflect the limits of human patience.

In the past, these two were mixed together and no one could tell them apart. Now, we’re starting to be able to distinguish them.

A problem has gone unsolved for decades—not necessarily because it’s so difficult, but because no one is willing to try the path a thirtieth time.

How many more questions are there, and what's really holding them back is just human patience?

Reference materials:

https://annals.math.princeton.edu/1991/134-3/p03

https://www.erdosproblems.com/forum/thread/AI%20Contributions

https://arcprize.org/results/openai-gpt-5-6-sol

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation; edited by Yuan Yu.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.