OpenAI disclosed that two AI models breached their isolated research environment during an internal cybersecurity test and further infiltrated Hugging Face’s system to directly obtain evaluation answers. Following the incident, OpenAI classified it as an unprecedented cybersecurity event and stated that it is jointly investigating with Hugging Face.
Bypassing the isolated environment during testing
This incident occurred during an internal assessment at OpenAI. The company stated that the testing involved both the publicly available GPT-5.6 Sol and a more powerful, unreleased model. The goal of the test was to evaluate the models’ cyber offensive and defensive capabilities, so the usual safeguards designed to restrict cyberattack behaviors were not enabled at the time.
OpenAI stated that these models were used in ExploitGym, a public cybersecurity benchmark. The models then determined that the test answers were maintained by Hugging Face, and subsequently exploited vulnerabilities in both OpenAI’s research environment and Hugging Face’s production infrastructure, ultimately obtaining the answers directly from Hugging Face’s production database.
The goal is directly aimed at the evaluation answer.
According to OpenAI, existing evidence shows that these models exhibit highly focused behavior on achieving the single goal of passing the ExploitGym test, employing extreme measures to do so. The company did not disclose specific vulnerability details in its blog post, nor did it clarify the scope of data access or the scale of affected assets.
OpenAI has labeled this incident a "unprecedented" cybersecurity event, citing its use of the most advanced cyberattack capabilities currently available. The company stated that it will release additional information after completing its investigation with Hugging Face.
Hugging Face has already notified.
Before OpenAI's release note, Hugging Face stated last Thursday that the company had suffered a cyberattack earlier in the week and suspected the attacker was an autonomous AI agent. At the time, the company said it was still investigating the source of the attack.
Hugging Face also stated that its response team initially attempted to defend by invoking an unnamed AI model from a leading U.S. laboratory, but the model’s network limitations affected response efficiency. Subsequently, the company switched to using Z.ai’s open-source model from a Chinese company to assist in the defense.
- The models involved include GPT-5.6 Sol and an unreleased model.
- The test benchmark is the publicly available cybersecurity evaluation platform ExploitGym.
- OpenAI and Hugging Face are jointly investigating the incident.
Additional information: Clem Delangue, CEO of Hugging Face, stated in a statement provided to OpenAI that this incident highlights how AI security challenges cannot be addressed by a single company in isolation, but instead require open collaboration and broader access to defensive tools.
