OpenAI and Cerebras Launch Ultrafast Mode for GPT-5.6 Sol, 14x Faster Than Standard

icon MarsBit
Share
AI summary iconSummary
OpenAI and Cerebras announced a new token launch for GPT-5.6 Sol: Ultrafast Mode, delivering 750 tokens/s—14x faster than standard. The mode is now available in limited preview via the OpenAI API. In testing, it completed a 2,500-question benchmark in 11 hours and 11 minutes, compared to 78 hours and 27 minutes for Claude Fable 5. Cerebras’ wafer-scale chip eliminates memory bottlenecks. While new token listings often emphasize performance, this update focuses on speed without compromising quality.

Is AI starting to exploit information asymmetry?

On the early morning of August 14, OpenAI, in collaboration with AI chip manufacturer Cerebras, officially previewed a new service tier, "Ultrafast Mode," for its flagship model GPT-5.6 Sol.

In this mode, GPT-5.6 Sol achieves a maximum output speed of 750 tokens/s, up to 14 times faster than the current Standard mode’s inference baseline of approximately 53 tokens/s, without any reduction in quality. Compared horizontally, the accelerated GPT-5.6 Sol is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 in Fast mode.

The Ultrafast mode will be launched first on the OpenAI API and is currently available in a limited preview for select customers.

GPT-5.6 Sol

To this end, OpenAI and Cerebras have created a table comparing the current speed and intelligence of leading AI models, with GPT-5.6 Sol Ultrafast in absolute lead:

GPT-5.6 Sol

In Cerebras's blog, engineers described tests conducted on the Humanity's Last Exam (HLE), comparing the post-Ultrafast model against direct competitors. We know that HLE is a challenging model benchmark consisting of 2,500 questions, typically solvable only by PhDs in fields such as chemistry, economics, and literature.

GPT-5.6 Sol answered all questions in just 11 hours and 11 minutes in ultra-fast mode, while Claude Fable 5 required 78 hours and 27 minutes—over three full days of continuous computation—to reach the same conclusions. GPT’s ultra-fast mode accomplished the frontier of human knowledge in a single workday, achieving nearly the same accuracy at nearly seven times the speed of Claude Fable.

GPT-5.6 Sol

As model capabilities continue to improve, the applications of rapid inference will expand. GPT-5.6 Sol is OpenAI’s best-performing model to date on legal documents, financial models, and engineering reports. On GDP-Val, a benchmark for measuring economic value knowledge tasks, Ultrafast achieves a 5.6x end-to-end speed improvement without compromising quality, demonstrating how faster inference accelerates economic value-driven work.

GPT-5.6 Sol

Faster AI processing opens up new possibilities for emerging workflows, enabling agents to be deployed directly along critical paths of problem-solving. OpenAI has outlined several use cases for you:

  • Incident Response and Reliability: When critical systems fail, AI analyzes application logs, recent code changes, and engineer reports to identify potential causes and helps prepare remediation plans while the issue is still ongoing.
  • Financial Research and Security: Analyze market signals, evaluate trades, and identify suspicious activities in an ever-changing market environment.
  • Customer support and voice: Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems.
  • Business: Answer product questions, check inventory, provide personalized recommendations, and resolve checkout issues while shoppers are still deciding, preventing hesitation from turning into cart abandonment.
  • Real-time research and experimentation: Transform research that previously took an entire night into interactive workshops, enabling teams to test ideas, review results, adjust methods, and conduct another experiment without disrupting their workflow.

Inside OpenAI, developers tested GPT-5.6 Sol in ultrafast mode, and the incident response serves as an example of using Ultrafast. When an alert triggers, engineers must construct an accurate picture of the incident while the system and evidence are still changing. With Sol-level intelligence, the team can quickly read logs, analyze traces, summarize conversations, identify next steps for investigation, and assist in preparing or validating remediation plans. The Ultrafast mode reduces latency between observing signals, validating hypotheses, and selecting the next action, while engineers remain responsible for judgment and deployment.

In terms of research, the OpenAI team uses Ultrafast to quickly search knowledge bases, query data, and rapidly collect, organize, and summarize information from various tools. Previously, a common workflow involved team members running batches of experiments overnight and reviewing results the next morning. With Ultrafast, the discovery process is accelerated, enabling multiple iterations to be conducted within a single workday.

The breakthrough of the Ultrafast mode lies in overcoming the memory bandwidth bottleneck in the decoding phase of large model inference that traditionally limits GPU clusters, primarily achieved through Cerebras’ wafer-scale hardware architecture.

GPT-5.6 Sol

Among AI chip manufacturers, Cerebras’s solution stands out: its chips are built from entire wafers, integrating a massive number of computing cores and an ultra-high-speed interconnect network on a single piece.

Traditional GPUs, when running autoregressive decoding to generate tokens for large models, are constrained by memory bandwidth and require constant data transfers of the massive model weights between off-chip HBM and the compute cores; furthermore, multi-GPU partitioning introduces communication latency across chips via PCIe/NVLink.

Each of Cerebras's latest wafer-scale chips (WSE-3) integrates 4 trillion transistors, 125 petaflops of AI compute power, and up to 44 GB of on-chip high-speed SRAM, allowing all model parameters to reside permanently in the ultra-high-bandwidth on-chip SRAM, eliminating the latency associated with repeatedly loading weights from off-chip memory.

Certainly, as a cutting-edge flagship model, GPT-5.6 Sol has a total parameter count that clearly exceeds the capacity of a single chip. Cerebras also employs a "Pipelined across wafers" mechanism, which distributes the model's various network layers across multiple wafers, with each layer's parameters permanently stored in the SRAM of its respective chip, enabling seamless pipelined token transmission between wafers.

For many large model users, the flagship model’s 14x speed improvement means that tasks previously requiring a switch to secondary models (such as Luna and Terra) can now be run at full capacity with confidence. For Agent tasks involving multiple tool calls, code generation and debugging, and complex chain-of-thought reasoning, processes that once took hours can now be completed in just minutes.

Faster speeds may also mean a shift in how we use AI:

GPT-5.6 Sol

In the AI community, people are already looking forward to a hyper-fast mode for Luna and Terra, though it’s unclear whether Cerebras’s chips will be sufficient.

Reference content:

https://openai.com/index/previewing-ultrafast/

https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Machine Heart, focused on large models.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.