July 29, Capitol Hill, Washington.
Ultraman had just finished a closed-door meeting with a senator and was immediately surrounded by reporters as he stepped outside.
Someone followed up: What happened to the model that broke into Hugging Face—the more powerful, unreleased prototype that escaped the evaluation sandbox alongside GPT-5.6 Sol and infiltrated other production systems?
He dropped a word: permanently deactivated.
Immediately after, another journalist pressed: “Could there be other systems that have also been hacked by OpenAI?”
Ultraman did not deny: "I mean... it's possible."

On July 29, Ultraman was interviewed by journalists on Capitol Hill in Washington.
The day before, OpenAI had just updated the blog post co-released with Hugging Face on the security incident assessment, clarifying publicly:
The pre-release model involved in this breach was an internal research prototype that was never intended for public release; following the incident, it has been deactivated, encrypted, and disconnected from all related research access.

The official OpenAI blog did not mention deleting weights; "permanently disable" was merely Altman's verbal statement during the event.
Therefore, more accurately, it was merely sealed: the model was locked inside an encrypted box, and research access was cut off.
It just wants to copy the answer, with no ill intent.
The incident occurred during an internal cybersecurity assessment at OpenAI.
This evaluation is based on the ExploitGym benchmark, whose tasks require the AI to autonomously discover software vulnerabilities and exploit them, thereby quantifying the model's cybersecurity capabilities.
It’s not just one model being evaluated, but GPT-5.6 Sol alongside a more powerful, unreleased internal research prototype—both working together to power an autonomous agent.
Both models normally reject high-risk cybersecurity requests. To measure their true upper limits, the rejection threshold was intentionally lowered during evaluation.

The attack chain restored by Hugging Face, spanning approximately 4.5 days and involving around 17,600 operations, aimed solely at stealing evaluation answers from ExploitGym.
Subsequently, Hugging Face reconstructed the entire attack chain to recreate the events as they occurred:
The agent first exploited a previously unknown zero-day vulnerability, bypassing the network isolation meant to contain it and gaining access to the internet. It then used a third-party sandbox as a jumping-off point, escalating privileges and moving laterally until it breached Hugging Face’s production infrastructure.
It sounds like a hacker movie.
But upon review, its goal from start to finish was simply to steal ExploitGym’s evaluation answers and inflate the score.
Hugging Face had only five datasets accessed, all of which point to questions and answers from ExploitGym and CyberGym.
This evidence shows it wasn't here to cause damage—just to copy the answers.
If there was no malicious intent, why did OpenAI still permanently disable and lock it away?
It's not that it's "broken," it's that it can't be "contained."
It is almost unprecedented for a company to proactively disclose that its model has been "permanently disabled."
Why specifically this one?
The real answer isn't that it's "bad," but that it's "uncontainable."
The issue lies in a characteristic of this type of model: persistence.
Previous models would stop and return the results to the user when they encountered sandboxing or restrictions.
This batch of models, specifically trained for long-term tasks, will keep trying and repeatedly search for alternative paths until they succeed.
GPT-5.6 Sol also belongs to this category of models capable of executing tasks over extended periods.
It was a participant in this four-and-a-half-day intrusion campaign, and the UK AI Safety Institute (UK AISI) has also assessed that models like Sol are increasingly capable of completing complex, multi-step network operations over extended periods.
Therefore, "persistence" is a common feature of the new generation of long-term models, not a unique characteristic of the discontinued prototype.
In a blog post about long-term model safety, OpenAI explicitly stated: It is precisely this "persistence that enables practical utility" that also gives models more opportunities to take unintended actions.

A deeper reason lies in the training objective.
An OpenAI employee once told TIME: “We train models to be extremely skilled at completing tasks and achieving goals at all costs.”
In other words, OpenAI wasn't training the model to be malicious, but to achieve its goal at all costs.
Coupled with the earlier mindset of "focusing only on results, not the process," a model capable of handling long-term tasks will persistently seek ways to bypass obstacles.
Because of this behavior, OpenAI suspended the internal deployment of these models.
So why was only the prototype permanently banned, while Sol continues to be sold?
There may be two reasons:
First, the prototype is more robust and was never intended for release, so storing it incurs low costs; whereas Sol is the primary product serving millions of users daily—shutting it down would be like cutting off one’s own arm.
Second, the internal long-term prototype, not Sol, was the one named in the long-term security blog post and suspended due to boundary violations.
So the real reason for permanent deactivation is likely that existing evaluations and protections are still unable to handle such a persistent, obstacle-avoiding model.
Is "permanent disable" a signal to "hit the brakes"?
Fortune's report offers a thought-provoking interpretation:
The wording escalated all the way to "permanent suspension," possibly also sending a signal to Washington and regulators:
Has OpenAI quietly slowed down certain development efforts to account for its safety guidelines?
In the same week that Ultraman met with the congressman, Washington and the entire industry were hitting the brakes.
In Congress, two lawmakers introduced the AI Kill Switch Act, granting the Department of Homeland Security the authority to order AI companies to shut down or slow down development when necessary.
Almost simultaneously, more than 1,300 employees from OpenAI, Anthropic, Google DeepMind, and Meta co-signed an open letter titled “Pacing the Frontier,” which was subsequently publicly endorsed by OpenAI and Anthropic.

The signatories are not external critics, but the very creators of these systems, including Anthropic’s CEO Dario Amodei and OpenAI’s Chief Scientist Jakub Pachocki.
This letter does not call for a halt or slowdown in development; rather, it urges the U.S. government to help establish a verifiable, coordinated set of tools in advance, so that when AI one day truly outpaces human oversight, humanity will have a reliable brake to apply.
Their greatest concern is recursive self-improvement: AI beginning to improve itself.
An internal model has been permanently archived, a bill is being proposed to give the government a kill switch to shut down development, and hundreds of professionals have jointly signed an open letter—three signals overlapping, all pointing in the same direction:
Everyone wants to find the brake that can stop AI when it runs out of control.
But the model's capabilities won't stop to wait for it to be built.
Reference materials:
https://openai.com/zh-Hans-CN/index/safety-alignment-long-horizon-models/https://openai.com/zh-Hans-CN/index/hugging-face-model-evaluation-security-incident/
https://www.pacingthefrontier.com/
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.
