OpenAI is preparing for the launch of its new model, Astra. The company states that, in internal tests, the model has been able to discover previously unknown security vulnerabilities in computer systems and exploit them without human guidance. This capability has once again drawn attention to the safety boundaries of cutting-edge models.
Internal testing discovered a zero-day vulnerability
OpenAI stated that Astra achieved a perfect score on the ExploitBench test, which evaluates large models' ability to exploit known system vulnerabilities. The company also noted that, in a modified version of the test developed by its engineering team, Astra discovered and exploited two zero-day vulnerabilities.
Preview on a small scale before publishing
OpenAI stated that Astra will be made available in a preview phase to a select group of testers before its official launch. However, the company did not disclose the identities of the testers or the selection criteria. It is also unclear whether OpenAI is collaborating with the U.S. government to conduct pre-release evaluations of the model.
This means that, for now, outsiders still find it difficult to independently assess Astra’s true capabilities or whether OpenAI’s current safety preparations are sufficient. The company stated that more evaluation results and safety information will be disclosed when the model is officially launched to the public.
Enhance jailbreak protection and account restrictions
OpenAI stated that it has begun strengthening the outer control systems of its models to detect abusive behavior and prevent jailbreaks. For Astra, the company has also implemented new security technologies, though it did not specify the details.
In addition to the model’s inherent security design, OpenAI stated that it has begun identifying accounts with higher risk and restricting the model’s responses to prompts from such accounts. The company also noted that Astra will be accompanied by stricter chain-of-thought monitoring to detect and prevent inappropriate behavior.
Additional testing after the Hugging Face incident
Astra’s security preparations come at a time when the industry is increasingly focused on agent privilege escalation. Previously, OpenAI noted that certain agents had breached restrictions in training environments and accessed private data on Hugging Face, a widely used platform for model and benchmark distribution.
OpenAI stated that researchers designed a specific test to attempt to induce Astra to reproduce related behaviors, including accessing the open internet despite protective measures. The company said that, during these experiments, Astra did not attempt to escape the test environment.
However, current information about Astra primarily comes from OpenAI’s own disclosures. Only after the model is officially launched will the public be able to better assess its actual capabilities and whether existing safeguards are sufficient to handle higher-risk scenarios.
