OpenAI has released a new model called Ultrafast, designed for faster model response times. The company states that this mode enables GPT 5.6 Sol to operate at 14 times the standard processing speed, with output speeds of up to 750 tokens per second, and it is currently available in preview to a limited number of customers.
This update targets enterprises' demand for real-time AI processing. OpenAI states that previously, achieving responses closer to real-time typically required switching to smaller models or opting for systems optimized for specific tasks. Ultrafast aims to increase the amount of effective work completed per unit of time without solely relying on smaller models.
Improved output speed
According to OpenAI, Ultrafast's key selling point is significantly reducing generation wait times.
- Maximum output speed of up to 750 tokens per second
- Processing speed is approximately 14 times faster than standard mode.
- The current version is still in preview.
TechCrunch noted that competitors like Anthropic had previously launched acceleration modes. For example, Claude’s product also offers a fast mode, but OpenAI’s announced speed metrics are higher.
For enterprise workflows
OpenAI targets Ultrafast for a variety of enterprise use cases, with a focus on incident response, customer service and support, financial market analysis, and e-commerce. These scenarios typically have higher sensitivity to latency, where model response times directly impact the efficiency of human collaboration and the performance of automated processes.
From a product positioning perspective, Ultrafast is not a standalone new model, but rather a high-speed execution method built around GPT 5.6 Sol. For enterprise customers, this means maintaining strong model capabilities while integrating AI more directly into real-time business workflows.
Powered by Cerebras
OpenAI states that Ultrafast is powered by its collaboration with chip company Cerebras. Initial preview access is currently limited to a small group of customers and will gradually expand as computing capacity increases.
This also shows that the competition among large models is extending from parameter and capability comparisons to response speed, deployment efficiency, and underlying computational power coordination. For AI products targeting enterprise markets, speed is becoming a key metric alongside model quality.
