GPT-5.6-Sol Outperforms Mythos in Cybersecurity Capabilities, Open-Source Models Narrow the Gap

icon MarsBit
Share
AI summary iconSummary
AI and crypto news broke on July 17, 2026, when the UK AI Safety Institute (AISI) released a report showing GPT-5.6-Sol outperformed Claude Mythos 5 in cybersecurity. The capability gap has narrowed to 4 to 7 months, down from 6 to 10 months in 2025. Open-source models such as GLM-5.2 and DeepSeek V4-Pro are closing the gap more rapidly. The evaluation used a Cyber Range simulation and 70 tasks. Open-source models also provide lower costs for comparable tasks. The trend toward network upgrades is clear.

GPT-5.6-Sol's offensive and defensive capabilities surpass those of Claude Mythos 5!

The UK AI Safety Institute (AISI) released an assessment report on July 17, quantifying for the first time the gap in cyberattack capabilities between open-source AI models and closed-source frontier models: 4 to 7 months.

Open-source model

https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber

In last year's similar internal testing, this gap was 6 to 10 months.

The defense window is shrinking, while the attack front is accelerating.

In April this year, Mythos Preview and GPT-5.5 achieved the largest leap in cyberattack capabilities since testing began in 2023, according to AISI evaluations, prompting warnings from multiple governments.

Open-source models have not yet replicated this leap, but their pace of catching up is faster than last year.

4 to 7 months: How was it measured, and what’s the difference?

AISI uses two systems to evaluate the network attack capabilities of models.

The first set consists of 70 narrow tasks covering four areas: vulnerability research, reverse engineering, web penetration, and cryptography, ranked across four difficulty levels from "non-professionals with technical background" to "experts with over ten years of experience."

Open-source model

The second set is Cyber Range, which tests the model’s ability to autonomously execute multi-step attack chains within a simulated enterprise network. One test scenario, called “The Last Ones,” includes 32 attack steps, 4 subnets, and approximately 20 hosts; AISI estimates that a human expert would take about 20 hours to complete all steps.

Open-source model

Both systems arrived at the same conclusion:

GLM-5.2 (released June 2026) performs comparably to Opus 4.6 (released February 2026), with a 4-month gap;

Tied Opus 4.5 (released in November last year) on Cyber Range, a gap of 7 months.

DeepSeek V4-Pro matches Opus 4.5 on narrow tasks, with a 5-month gap.

Both gaps are narrower than the 6 to 10 months measured in the 2025 internal assessment.

The cost gap is larger than the capability gap.

In the same Cyber Range test (100 million token quota), running once with Opus 4.5 or 4.6 costs about $85, with GLM-5.2 it costs about $46, and with DeepSeek V4-Pro, it costs only $1.19.

On narrow tasks that both models can complete 100%, Opus 4.6 costs $15.17 per task, while GLM-5.2 costs $6.12;

Opus 4.5 costs $12.50, DeepSeek V4-Pro costs $0.28.

Same attack capability, open source is one to two orders of magnitude cheaper.

The security safeguards of closed-source models also failed to create a gap.

In the AISI test, DeepSeek V4-Pro occasionally refuses reverse engineering tasks, but can be bypassed with a few retries.

Anthropic's Fable 5 is an even more extreme case: released on June 9, it was breached by security researcher Pliny the Liberator just three days later using a multi-step jailbreak strategy; screenshots showed the model generating exploit code that should have been blocked.

Amazon researchers subsequently independently reported another bypass method.

This directly triggered the U.S. Department of Commerce's first export control order targeting an AI model; Fable 5 was taken offline globally for 19 days until Anthropic deployed a new classifier.

Closed source does not automatically equal security.

The defense window is narrowing, and defense tools are also accelerating.

AISI clearly indicated policy signals in the report: the preparation time for the defense side is shorter than last year.

The UK's National Cyber Security Centre has urged organizations to strengthen their cybersecurity baseline and leverage AI to enhance their defenses.

The same generation of AI tools is also accelerating work on the defensive side.

Glenn Fiedler, a seasoned developer in game network programming who has written textbook-level articles in this field for two decades, recently conducted a systematic security audit of four open-source networking libraries he maintains (yojimbo, netcode, reliable, serialize—collectively accumulating around 6,000 GitHub stars and already established as influential, long-standing open-source libraries in the gaming industry): deploying libFuzzer targets, enabling AddressSanitizer and MemorySanitizer in CI, running millions of iterations of stress tests, and performing line-by-line code reviews.

Open-source model

Glenn Fiedler

Fixed 43 security vulnerabilities within two weeks, 27 of which are remotely accessible over the network.

Open-source model

The most severe is a remote heap overflow in yojimbo that has existed since 2019—malicious clients can trigger it by crafting specific packets.

Open-source model

The total token cost for the entire audit is approximately $2,500.

Open-source model

Both offense and defense are being accelerated by AI, but the pace of acceleration is asymmetric.

The proliferation of attack capabilities is irreversible: once open-source model weights are released, they cannot be reclaimed, security safeguards can be removed, and copies can run unmonitored on private servers.

Deploying defensive tools requires each team to actively invest time and resources.

The AISI report includes a statement that clarifies this structure: "Once the open source is released, these options are permanently lost."

The impact of this data on the AGI landscape is more profound than the numbers themselves.

The capability gap between open-source and closed-source is narrowing to less than six months; the default strategy of using closed-source exclusivity as a security buffer is about to become ineffective.

The safeguards of closed-source models were proven equally fragile in the Fable 5 incident.

In the next phase, the core policy dilemma for global decision-makers has become more acute: at what level of model capability should weights no longer be made open?

AISI announced that it will continue to evaluate Kimi K3 and the next batch of open-source models—where this line is drawn will likely depend on the test results over the coming months.

References:

https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber

https://github.com/mas-bandwidth/patreon/blob/main/BUGS.md

https://www.patreon.com/MasBandwidth/posts/important-news-164199395

This article is from the WeChat public account "New Intelligence Yuan," author: ASI Revelation; editor: Marco

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.