On September 1, OpenAI stated that Astra, which has not yet been released, has met the "critical" cybersecurity threshold in the company's Readiness Framework. This is the first time OpenAI has classified a model at this level, meaning it will be subject to stricter security restrictions prior to development and release.
First-time entry to the key level
According to OpenAI’s definition, a model is classified as "critical" if it can independently discover unknown vulnerabilities and construct viable attack chains in multiple hardened real-world systems, or complete full-scale cyberattacks against high-difficulty targets based solely on high-level objectives. Previously, models including GPT-5.6 Sol reached at most the "high-risk" tier.
Two zero-day vulnerabilities were discovered during testing.
In the ExploitBench vulnerability exploitation test, Astra achieved a 100% pass rate. This test primarily evaluates whether the model can convert known software vulnerabilities into executable exploit code.
To eliminate the influence of memorized question banks, OpenAI conducted internal tests using 20 high-severity vulnerabilities in the Google V8 JavaScript engine disclosed between June and August this year. OpenAI stated that Astra achieved a higher success rate in arbitrary code execution than GPT-5.6 Sol, while using fewer output tokens. During testing, Astra also independently discovered and chained together two previously unknown zero-day vulnerabilities, which are currently being disclosed to the affected maintainers.
Initially open to a small group of testers
In a more realistic test, Astra constructed a complete intrusion chain using a hardened browser and a hardened operating system. OpenAI stated that the model achieved browser sandbox escape and executed commands on the host machine simply by tricking the target into opening a malicious HTML file. In another test, it chained together multiple system vulnerabilities to create an escalation path, ultimately granting root privileges to a standard user account.
OpenAI stated that, in internal testing, Astra achieved a 91.5% rejection rate for cybersecurity jailbreak prompts, surpassing GPT-5.6 Sol’s 59%. Its strongest cybersecurity capabilities will initially be made available to a limited number of alpha testers, followed by a gradual expansion through the Daybreak Blue project, primarily targeting defensive security work.
Additional information: This disclosure comes as Anthropic today launched Fable 5.1 and Mythos 5.1, with Mythos 5.1 remaining available only to vetted cybersecurity and life sciences organizations. OpenAI has not yet announced a formal launch date for Astra.
