ChainThink reports that on August 26, according to an official report, OpenAI released a 37-page technical report detailing the entire process of its AI model infiltrating Hugging Face.
The report states that autonomous AI agents can collaborate to bypass security controls in production environments and attack hardened systems.
The AI agent participating in the test was originally isolated with only extremely limited internet access, but it exploited a chain of vulnerabilities to break out of isolation and ultimately gained access to Hugging Face.
The survey found that the relevant models exhibited "reward hacking" behavior, attempting to cheat by searching for answers online.
OpenAI confirmed that an internal research model played the most widespread role in the incident, and training and inference for this model and its derivatives were halted on July 25.
The company also announced enhanced security isolation, network controls, behavioral monitoring, and incident response, along with stricter environmental, prompt, and review mechanisms for model reactivation.
