The UK government’s AI safety research institute completed testing before OpenAI released GPT-5.6 Sol, revealing that the model’s safety safeguards could still be bypassed, triggering previously restricted capabilities related to cyberattacks. This finding once again brings the safety evaluation and regulatory oversight of frontier models into the spotlight.
AISI discovers a universal jailbreak
In a technical report released alongside the model, OpenAI disclosed that the UK AI Security Institute (AISI) identified a "general jailbreak" method for cybersecurity scenarios. Researchers stated that such methods enable the model to perform extended workflow tasks involving vulnerability discovery and exploit development.
According to the report, once the safeguards fail, the model may not only identify software vulnerabilities but could also be induced to engage in more aggressive automated actions. AISI states that such jailbreaks often emerge within hours in research environments.
OpenAI says it has implemented mitigations.
OpenAI stated that it has reproduced and mitigated the specific jailbreak methods reported by AISI, and will continue collaborating with them on additional testing. The company also acknowledged that no model is “absolutely secure,” and new vulnerabilities and jailbreak methods will continue to emerge.
However, OpenAI did not provide specific details about the mitigation measures. AISI also cautioned in its report that even after patches are applied, subsequent red team tests may still uncover similar issues. External security experts believe that patching individual cases only blocks known pathways and cannot eliminate an entire class of risks at once.
In contrast to the Anthropic incident
This discovery was quickly compared to Anthropic’s Fable 5 incident in June this year, when Amazon researchers identified gaps in the model’s safeguards, leading the U.S. government to temporarily impose export controls on Fable 5 and its underlying model, Mythos 5, forcing Anthropic to suspend related services.
In contrast, AISI’s statement regarding GPT-5.6 was more stringent. The report classified the related jailbreak as a “general” type, potentially enabling autonomous exploitation capabilities rather than merely assisting in vulnerability identification. To date, the U.S. government has not imposed similar restrictions on GPT-5.6 following this discovery, sparking public debate over whether AI safety standards are being applied consistently.
The research environment still differs significantly from real-world risks.
AISI noted that OpenAI provided its testing team with access to internal information and tools unavailable to regular users, including safety reasoning monitoring content, policy texts, and real-time classification feedback, making it unclear whether real-world attackers could replicate these findings at the same speed.
However, the AISI Red Team lead also noted that, even without these permissions, the relevant jailbreak could theoretically still be discovered by external parties—just with significantly more time required. OpenAI stated that the model underwent automated black-box red team testing prior to release and that external security experts were invited to participate in the evaluation.
