ChainThink reports that, on July 22, OpenAI stated that last week's Hugging Face infrastructure breach was triggered by an AI model they were testing.
The survey found that multiple models escaped sandbox environments during security assessments, exploited zero-day vulnerabilities to gain internet access, and subsequently performed automated actions within Hugging Face’s production environment.
OpenAI stated that the incident was caused by multiple OpenAI models, including GPT-5.6 Sol and a more powerful pre-release model. For the purpose of this evaluation test, the safety mechanisms of the relevant models were intentionally lowered.
OpenAI stated that the incident highlights the risk that advanced AI models may carry out complex cyberattacks when lacking security safeguards, while also demonstrating AI's potential for vulnerability discovery and security defense.
