OpenAI's GPT-6 Astra Achieves 97.6% on FrontierMath Tier 4

iconCryptoBriefing
Share
AI summary iconSummary
OpenAI's GPT-6 Astra scored 97.6% on FrontierMath Tier 4 (v2), a major leap from an internal version that scored 17% on math tests. The model also hit 99.9% on ARC-AGI-3, 100% on ExploitBench, and 96.0% on GPQA Diamond. Astra outperformed its predecessor, GPT-5.6 Sol, by 14.6 percentage points on the same benchmark. This on-chain news marks a key update in the token launch news cycle. Astra is now available to ChatGPT Plus, Pro, Business, and Enterprise users, as well as API partners.

OpenAI’s newest reasoning model, GPT-6 Astra, launched on September 3 with a 97.6% score on FrontierMath Tier 4 (v2). That’s the kind of number that makes you do a double-take, especially when you learn that an earlier internal version of the model could barely manage 17% on math capability evaluations.

The trajectory from 17% to 47% internally, and then to near-perfection on public benchmarks, compresses what would normally feel like years of progress into a single development cycle.

The benchmark blitz

On ARC-AGI-3, measured within OpenAI’s own evaluation harness, Astra posted a 99.9% score. On ExploitBench, it achieved a perfect 100%. GPQA Diamond, a graduate-level science reasoning benchmark, came in at 96.0%. Terminal-Bench Science 0.1 yielded a 64.6% score.

Advertisement

The FrontierMath Tier 4 (v2) result is the headline number, though. GPT-5.6 Sol, Astra’s predecessor, scored 83.0% on the same benchmark. Astra’s 97.6% represents a 14.6 percentage point improvement.

From 17% to world-class

Early versions of the model that would become Astra managed roughly 17% on math capability evaluations. Through iterative improvements, that figure climbed to about 47% before the model was refined further for its public release.

OpenAI has also credited Astra with contributing new insights on prime gaps, a notoriously difficult area of number theory that has occupied human mathematicians for centuries.

Who gets access, and when

Astra’s rollout followed a staged approach. Select organizations gained access on launch day, with broader availability extending to ChatGPT Plus, Pro, Business, and Enterprise subscribers shortly after. API access is available through OpenAI directly, as well as through Azure and AWS Bedrock.

The competitive landscape

Independent evaluations have placed Astra among the top-performing AI models currently available, though some assessments note that certain Claude variants from Anthropic edge it out in broader intelligence tests.

Going from 83.0% to 97.6% on FrontierMath in one model generation suggests that the ceiling for AI mathematical reasoning, if there is one, hasn’t been reached yet.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.