Just now, the final FrontierMath Tier 4 challenge was solved by GPT-6 Astra.
At this point, all questions in this research-level test, once considered a mathematical nightmare for large models, have been successfully solved by AI at least once.
Epoch AI has also officially concluded that FrontierMath Tier 4 is saturated.

Although the chart above shows that Astra has not yet directly reached 100%, OpenAI's own reported result is 97.6%.
However, according to Epoch's algorithm, "solved all" means that, after aggregating attempts across different models and time periods, each question in Tier 4 has been successfully solved at least once.
And what Astra took down was precisely the last one that had not yet been breached by AI.
Simply put, Astra may have made a mistake during one of the tests, but the problem it got right was one that hadn't been solved before.

To be honest, this rate of development is still insane.
If you follow model evaluations, you know that when Tier 4 was first launched on July 11, 2025, the highest score on the leaderboard was only around 5%.
Just over a year and two months have passed, and this once formidable "wall" has become a "stepping stone"—the final challenge that AI could not easily bypass has now been overcome.
The problem setter, Jay Pantone, associate professor of mathematics at Marquette University, also noted that previously, AI would always seek numerical shortcuts, but this time, the AI’s solution was very close to his own.

However, with GPT-6 Astra recently launching a fierce assault on the mathematical community, solving Millennium Prize problems left and right, Pantone says they can hardly be surprised anymore!
Mathematics has arguably become one of the areas where Astra has most prominently increased its visibility recently.
From less than 2% to adding a dedicated "research-grade defense."
FrontierMath was first released on November 7, 2024, with the sole purpose of preventing mathematical benchmarks from being rapidly exploited by AI.
At the time, traditional math benchmarks like GSM8K and MATH were becoming increasingly ineffective at distinguishing between top models, so Epoch AI collaborated with over 60 mathematicians to create a new set of previously unpublished original problems.
Among them are Fields Medalists such as Terence Tao, Timothy Gowers, and Richard Borcherds.
After reviewing some of the research-level problems, Terence Tao bluntly stated that these questions were “extremely difficult,” and even suggested that the hardest Tier 3 problems might continue to stump AI for several more years.

After the first round of testing, as expected, the leading model's accuracy was less than 2%.
Initially, FrontierMath’s core question bank consists of 300 questions, categorized into three difficulty levels: Tier 1, Tier 2, and Tier 3.

Tier 1 is roughly equivalent to advanced undergraduate problems and math Olympiad questions, but allows the use of more advanced tools; Tier 2 reaches the level of advanced graduate studies; Tier 3 is closer to the exploratory research problems encountered by early-stage PhD students.
But as reasoning models emerged, these first three tiers became increasingly insufficient, so in 2025, Epoch added another layer: Tier 4.
Tier 4 problems are primarily designed by mathematics professors and postdoctoral researchers, each conducting a short-term study lasting several weeks around their specific research focus, ultimately condensing their findings into a single problem that can be automatically verified.

Initially, Tier 4 included 50 questions, covering areas such as analysis, number theory, combinatorics, topology, and algebraic geometry.
At launch, across all model tests combined, only three questions had ever been solved, and those solutions relied on certain correct but insufficiently justified assumptions.
Given this, Epoch once also stated on the official website's sample questions page: Some of these questions may not be solved by AI for decades.

(This statement is now being repeatedly beaten to death)
However, as models have become more powerful, another issue has emerged: the question bank itself is beginning to fail under AI scrutiny.
During testing, OpenAI found more errors in FrontierMath than expected.
Subsequently, Epoch initiated an independent audit, first using GPT-5.5 and Claude Opus 4.7 to screen for potential issues, then assigning mathematicians to review each question individually. The final v2 version was released in June 2026, with Tier 4 correcting 12 questions, removing 7, and retaining 43.
Even after revisions, the model's performance continues to rise.
GPT-5.6 Sol achieved 83.0%, Claude Fable 5 reached 90.2%, and GPT-6 Astra has now attained 97.6%.

More importantly, Astra also solved the one problem that had never before been solved by AI.

At this point, every question in Tier 4 has been successfully solved by AI at least once. This research-level defense, specifically designed for the most powerful models, has ultimately been breached.
However, this does not mean that "math has been solved by AI." On the contrary, FrontierMath has already begun moving to the next stage.
In addition to Tiers 1–4, the project now includes genuine Open Problems and the formalization of Erdős open problems in Lean through FrontierMath Erdős.

The former tests the model with research problems that have not yet been solved by the mathematical community; the latter requires the AI to generate complete proofs that can be formally verified.
This time, Astra solved only 2 out of FrontierMath Erdős's 68 problems.
The story continues~
This article is from the WeChat public account "Quantum Bit," authored by Henry.
