OpenAI’s newest reasoning model, GPT-6 Astra, launched on September 3 with a 97.6% score on FrontierMath Tier 4 (v2). That’s the kind of number that makes you do a double-take, especially when you learn that an earlier internal version of the model could barely manage 17% on math capability evaluations.
The trajectory from 17% to 47% internally, and then to near-perfection on public benchmarks, compresses what would normally feel like years of progress into a single development cycle.
The benchmark blitz
On ARC-AGI-3, measured within OpenAI’s own evaluation harness, Astra posted a 99.9% score. On ExploitBench, it achieved a perfect 100%. GPQA Diamond, a graduate-level science reasoning benchmark, came in at 96.0%. Terminal-Bench Science 0.1 yielded a 64.6% score.
The FrontierMath Tier 4 (v2) result is the headline number, though. GPT-5.6 Sol, Astra’s predecessor, scored 83.0% on the same benchmark. Astra’s 97.6% represents a 14.6 percentage point improvement.
From 17% to world-class
Early versions of the model that would become Astra managed roughly 17% on math capability evaluations. Through iterative improvements, that figure climbed to about 47% before the model was refined further for its public release.
OpenAI has also credited Astra with contributing new insights on prime gaps, a notoriously difficult area of number theory that has occupied human mathematicians for centuries.
Who gets access, and when
Astra’s rollout followed a staged approach. Select organizations gained access on launch day, with broader availability extending to ChatGPT Plus, Pro, Business, and Enterprise subscribers shortly after. API access is available through OpenAI directly, as well as through Azure and AWS Bedrock.
The competitive landscape
Independent evaluations have placed Astra among the top-performing AI models currently available, though some assessments note that certain Claude variants from Anthropic edge it out in broader intelligence tests.
Going from 83.0% to 97.6% on FrontierMath in one model generation suggests that the ceiling for AI mathematical reasoning, if there is one, hasn’t been reached yet.
