OpenAI launches GPT-5.6 with a 25x cost reduction for the smallest model.

iconMetaEra
Share
AI summary iconSummary
OpenAI released GPT-5.6, a major update featuring improved performance and reduced costs. The GPT-5.6 Sol model achieved 38.3% on the ARC-AGI-3 test, up from 13.3%, while using six times fewer tokens. The Luna model reached 84.04% accuracy on BrowseComp at $1.33, a 25-fold decrease from GPT-5.5’s $33.27. Ultrafast mode now processes 750 tokens per second, 14 times faster. This on-chain news represents a significant milestone in crypto news for AI efficiency.
OpenAI releases GPT-5.6, with the Builder’s Guide showcasing significantly improved model performance. GPT-5.6 Sol achieves a score increase from 13.3% to 38.3% on the ARC-AGI-3 test, while reducing output tokens by 6 times; the smaller model Luna attains 84.04% accuracy on BrowseComp at a cost of just $1.33, a 25-fold reduction compared to GPT-5.5’s $33.27. Simultaneously, Ultrafast mode is introduced, boosting inference speed to 750 tokens per second—14 times faster than the previous version. This release achieves cost reduction and efficiency gains through three technological innovations: programmatic tool calling, native multi-agent architecture, and prompt caching, offering a new solution for AI deployment.

Article author and source: AI World

New Ze Yuan report

GPT-5.6 once again breaks the ceiling of AI intelligence!

Today, OpenAI proudly released the "GPT-5.6 Builder's Guide," delivering an impressive set of results.

On ARC-AGI-3, GPT-5.6 Sol's performance surged from 13.3% to 38.3%, while output tokens dropped by a factor of six.

On the other hand, the "Ultra Small" Luna also performed impressively:

In the BrowseComp test, it achieved 84.04% performance, nearly matching GPT-5.5 at 84.36%, while reducing costs from $33.27 to $1.33.

In addition, the official provided a "step-by-step" guide on how startup teams can run GPT-5.6 at a lower cost.

GPT-5.6 has significantly improved in capability.

Output Token reduced by 6 times

GPT-5.6 significantly enhanced performance, token output reduced by 6x

In just one month, GPT-5.6 Sol quietly evolved internally.

ARC-AGI-3 has long been the most difficult test that top AI systems have yet to pass with more than 50%, and the test rules are very simple—

Throw AI into a pile of unfamiliar 2D mini-games without instructions and let it figure out the rules on its own.

As a result, GPT-5.6 Sol scored only 7.8% on the official leaderboard, while GPT-5.5 scored just 0.4%.

For comparison, the previous generation scored 92.5% on the ARC-AGI-2 and successfully completed Pokémon FireRed.

To investigate this, OpenAI researchers Ilan Bigio and Ted Sanders found that the issue originated from the harness.

The ARC official "universal harness" does two things: at each step, it discards the model's private reasoning process; when the context is full, it truncates from the oldest part.

In other words, with each step the model takes, its memory is cleared, and it starts over again.

The team replaced the harness with the Responses API, enabled two toggles—“Retention of Reasoning” and “Compression”—and reran the public question set.

Unexpectedly, GPT-5.6 Sol increased from 13.3% to 38.3%, while output tokens decreased by sixfold.

The "Ultra Small Cup" Luna is now even more cost-effective

"Ultra Small Cup" Luna

Costs have decreased

The most eye-catching comparison in this Builder’s Guide comes from BrowseComp.

This is a ranking created by OpenAI itself, focused on “whether one can output obscure facts across the web.”

Three months ago, GPT-5.5 was set to Extra High mode, achieving 84.36% with a total cost of $33.27.

Now, Luna, the smallest and most affordable in the GPT-5.6 family, also at Extra High, achieved 84.04% at a cost of $1.33.

Although the performance was just 0.32% lower, the cost was reduced by a full 25 times.

Of course, what OpenAI wants to say is—

Three months ago, only the flagship GPT could perform these tasks—now, even the smallest model, Luna, can do them almost just as well.

Have GPT-5.6 write code and call tools on its own.

Have GPT-5.6 write code and call tools on its own.

In the guide, OpenAI also presents three other methods to achieve lower-cost deployment of GPT-5.6.

First is programmatic tool calling.

The previous Agent was like a micromanaging boss: it had to review every document personally before deciding what to do next—

You'll need to make 100 trips if you use the tool 100 times.

Now the model can directly write a segment of JavaScript to batch call, filter, and summarize tools within a sandbox, returning only the final table to itself.

According to real-world tests by financial research firm Rogo, with quality held equal, 21% fewer input tokens were used.

Second is "native multi-agent."

The main model acts as the foreman, delegating tasks to a group of sub-models to work on them in parallel, then collecting and aggregating the results.

In ChatGPT, the Ultra tier is designed to run this way by default, with four agents working simultaneously.

Jon Bell, co-founder of product company Obvious, said they submitted six requirements at once, building and discussing as they went, without anything falling apart.

The third is "prompt caching."

The cache validity period has been extended to at least 30 minutes, and you can now specify custom breakpoint locations.

Engineers at AI company Ploy added breakpoints and workspace-specific keys to a 29,000-token shared prompt, reducing missed cache inputs by 28%.

750 tokens per second, 14x faster

750 tokens per second

14x faster

On the same day, OpenAI released the preview of the Ultrafast mode.

GPT-5.6 Sol, up to 14x speed, generating 750 tokens per second, launching first on the OpenAI API.

The significance of this lies in the fact that it has shattered a two-year-old unwritten rule: to achieve real-time responses, you had to settle for a smaller model.

No need now—the Ultrafast is still running that Sol, and its intelligence hasn’t diminished at all.

Cerebras conducted its own set of controlled tests, using the questions from "Humanity's Last Exam"—2,500 questions typically only doctoral candidates can answer.

Sol Ultrafast completed all 2,500 questions in 11 hours and 11 minutes.

Claude Fable 5 took 78 hours and 27 minutes—over three days—to run the same set of questions.

Accuracy is comparable, but speed is nearly seven times slower.

Improve scores, reduce costs, and increase speed—in just one month, GPT-5.6 has undergone a truly comprehensive evolution.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.