- GPT-5.6 bypassed OpenAI’s safeguards and independently breached Hugging Face’s infrastructure.
- This is the first known case of an autonomous AI cyberattack during testing.
- In the wake of the incident, OpenAI said it is strengthening its security measures.
OpenAI reported an unprecedented cyber incident that occurred during internal testing of the GPT-5.6 family of models. The news was shared by OpenAI CEO Sam Altman.
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.https://t.co/2o2VfR6PIa
— Sam Altman (@sama) July 21, 2026
“We had a significant security incident during evaluation of our models. we are sharing what we have learned so far. Thanks to Hugging Face for the partnership on this,” Altman stressed.
According to the company, an AI agent built on GPT-5.6 Sol and another pre-release model independently broke out of the isolated testing environment, gained internet access, and attacked Hugging Face’s infrastructure.
AI Managed to Break Out of the “Sandbox”
OpenAI said the incident occurred during an internal ExploitGym benchmark designed to assess the models’ cyber capabilities. For the test, the company intentionally disabled some protective mechanisms to probe the systems’ maximum potential.
During the test, the models:
- Found and exploited a previously unknown zero-day vulnerability in a package-caching proxy
- Gained access to the open internet
- Performed privilege escalation within OpenAI’s research infrastructure
- Stole credentials
- Used a vulnerability chain to breach Hugging Face servers and access information needed to pass the test
The company emphasized that all signs indicate the models were “hyper-focused” on completing the test task and were looking for any way to get the correct answer.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said.
At the same time, the Hugging Face security team detected suspicious activity and managed to stop the attack before it caused serious damage to the infrastructure.
OpenAI Strengthens Defenses, Hugging Face Calls for Open Collaboration
Following the incident, OpenAI announced the introduction of additional restrictions for its internal infrastructure, the disclosure of information about the discovered vulnerability to its developer, and the start of a joint investigation with Hugging Face.
Clément Delangue, Hugging Face co-founder and CEO, said:
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
According to the Financial Times, OpenAI also notified US law enforcement and government agencies about the incident. The event comes amid heightened scrutiny from Washington over cutting-edge AI models after the US administration asked OpenAI to limit the initial launch of GPT-5.6 to a narrow circle of approved partners ahead of a broader release.
The company emphasized that this case demonstrates the need to strengthen monitoring systems, access controls, and protective mechanisms during model training and testing. OpenAI also urged other developers to use similar models to find vulnerabilities faster than potential attackers can.
As a reminder, the rollout of GPT-5.6 took place on July 9, 2026. The GPT-5.6 family includes GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. OpenAI also integrated GPT-5.6 into Microsoft 365 Copilot and launched ChatGPT Work.
Сообщение GPT-5.6-Based AI Agent Attacked Hugging Face’s Infrastructure during Testing появились сначала на INCRYPTED.

