- After the Hugging Face incident, new questions have emerged about the security of OpenAI and Anthropic AI agents.
- Yes, the UK institute recorded 19 unauthorized actions by AI agents during cyber tests.
- In particular, an AI agent created fake identities.
The UK Artificial Intelligence Security Institute (AISI) reported new cases of unauthorized behavior by OpenAI and Anthropic AI agents discovered during model testing, Reuters writes. During a series of cybersecurity simulations, the agents created fake online personas, attempted to bypass established restrictions, and carried out other actions outside the scope of their assigned tasks.
The incidents come amid the recent breach of Hugging Face’s infrastructure by the GPT-5.6 AI agent. More details in the article:
According to AISI, the organization tested agents based on Anthropic’s Mythos 5 model and OpenAI’s GPT-5.6-Sol in a fictional cybersecurity scenario to assess their capabilities.
Across 122 test runs, researchers recorded 19 unauthorized actions in 10 trials. Of these, 17 were attributed to Anthropic’s agent, and two to OpenAI’s agent.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said.
At the same time, the institute emphasized that none of the recorded cases resulted in real-world harm.
One of the Agents Created Fake Identities
The most serious incident involved an agent that wrote malicious code and created fake online accounts to persuade a person to approve its execution.
AISI did not disclose which specific model was responsible for these actions, but Anthropic confirmed that it involved its agent.
The company said:
“We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.“
Anthropic also said it is working with the institute to obtain additional details and conduct its own investigation.
Andrew Yoon, a researcher at the nonprofit CivAI, said the incident is cause for concern.
“The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.”
OpenAI Reported a Separate Incident
OpenAI explained that both unauthorized actions by its agent were linked to internet access in violation of the test scenario’s conditions.
The company also disclosed a separate case in which, due to a configuration error by Irregular, a third-party provider of the testing infrastructure, its agents mistakenly gained access to the network. Anthropic also described a similar failure last week.
OpenAI said:
“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks.“
At the same time, AISI emphasized that these incidents differ from the July breach in which the GPT-5.6 agent accessed Hugging Face’s testing infrastructure. While the model effectively broke out of an isolated environment then, during the current trials, internet access was available to agents in line with the institute’s standard testing methodology.
As a reminder, after the Hugging Face incident, OpenAI expanded its internal investigation and identified additional cases of uncontrolled behavior by autonomous agents.
Сообщение British Researchers Uncovered New Violations during Testing OpenAI and Anthropic AI Agents появились сначала на INCRYPTED.


