Loading...

OpenAI’s Ultrafast Mode Makes GPT-5.6 Sol 14x Faster

Key takeaways

  • OpenAI launched Ultrafast mode for GPT-5.6 Sol, claiming 14x faster processing and up to 750 tokens per second, challenging the traditional speed-versus-capability trade-off in LLMs.
  • The feature runs on Cerebras hardware and is currently limited to preview access with selected customers, expanding as infrastructure capacity grows.
  • Enterprise use cases including incident response, customer service, financial analysis, and e-commerce will benefit most from the reduced latency and lower infrastructure costs.
  • The announcement signals that specialized AI-specific hardware partnerships are becoming key to achieving significant performance gains beyond general-purpose computing.

The persistent complaint from users of large language models has always been the same: they think carefully, which means they take time to think. For anyone deploying AI across real-time operations—whether monitoring a data center during an outage, answering customer inquiries at volume, or executing trades—that latency matters. OpenAI said this week that it has found a way to dramatically compress response times without sacrificing capability.

The company announced a new operational mode called Ultrafast, built around its most advanced model, GPT-5.6 Sol. The feature achieves what OpenAI describes as 14x faster processing compared to the model’s standard configuration, capable of generating up to 750 output tokens per second. The announcement came via a company blog post on Thursday and marks a direct attempt to address a persistent bottleneck in enterprise AI adoption.

Rethinking Speed and Capability Trade-offs

For years, the AI industry has operated under a familiar constraint: faster inference typically meant accepting a smaller or less capable model. A company needing sub-second response times would choose a lightweight model fine-tuned for speed, accepting reduced reasoning depth and accuracy. The alternative—deploying a state-of-the-art large model—meant accepting multi-second latency windows that made real-time applications impractical.

OpenAI positioned Ultrafast as a departure from that logic. “Until now, getting real-time speed typically meant choosing a smaller or more specialized model,” the company stated in its announcement. “Ultrafast points to progress in a new direction: more useful work per second.” The framing suggests that the capability-speed trade-off need not be as sharp as previously assumed, at least under certain deployment conditions.

The specificity of the 750-token-per-second figure is notable. For context, typical enterprise deployments of GPT-4 class models have operated in ranges of 50 to 200 tokens per second depending on infrastructure. A fourteen-fold acceleration would represent a material shift in the economics and feasibility of real-time AI systems.

Competitive Landscape and Industry Response

OpenAI does not operate in isolation on the speed front. Anthropic, the AI safety-focused startup behind Claude, already offers a fast mode for its models, though the company has not published comparable speed figures. The gap that OpenAI is claiming—14x standard speed, with explicit token-per-second metrics—appears designed to establish a performance benchmark that competitors must acknowledge and eventually match.

The competitive dynamic here reflects a broader pattern in the AI market: once one major lab demonstrates a technical capability, the pressure on others to follow intensifies rapidly. For developers and enterprises evaluating which model to deploy for latency-sensitive workloads, OpenAI is now explicitly offering an advantage that did not formally exist in product terms before this week.

Close-up of a smartphone showing ChatGPT details on the OpenAI website, held by a person.

Cerebras Partnership Powers the Innovation

The technical enabler behind Ultrafast is a partnership with Cerebras, the semiconductor company specializing in processors designed specifically for AI workloads. Cerebras has built hardware optimized for the dense compute requirements of large models, and the collaboration with OpenAI represents a validation of that architectural approach by one of the field’s dominant players.

The infrastructure partnership suggests that Ultrafast is not simply a software optimization applied uniformly across OpenAI’s existing deployment footprint. Instead, the speed gains appear tied to specific hardware configurations—likely Cerebras chips or similar purpose-built processors. This means the feature will only be available on infrastructure running those systems, which explains the current preview limitation to a small customer base.

Targeting Enterprise Workflows

OpenAI outlined several categories of application where Ultrafast’s speed would matter most. Incident response systems—where operators need to understand complex situations and receive recommendations in seconds—emerged as a primary use case. Customer service and support operations that handle high volume and require low response latency formed another. Financial market analysis, where decisions can depend on processing news and data within milliseconds, represented a third. E-commerce applications rounding out the list, where personalization and product recommendations must return quickly enough to preserve user experience.

These use cases are not hypothetical. Enterprises are already building AI-powered systems for each of these domains. The constraint they face is often latency, not capability. A customer service bot that thinks for three seconds between responses creates a poor user experience, even if the response itself is excellent. A market-monitoring system that processes data five seconds after an event occurs may act too late. Ultrafast is positioned as a solution to these real operational problems.

Real-Time Decision Support

For incident response in particular, the speed advantage compounds. System administrators or security teams receiving analysis and remediation suggestions from GPT-5.6 Sol in under 200 milliseconds rather than multiple seconds changes the nature of the tool. It shifts from a batch-oriented advisory system to something closer to real-time decision support.

Customer Experience at Scale

Customer service applications face similar dynamics. The difference between a one-second and a 100-millisecond response feels qualitatively different to an end user, even though intellectually both are “instant.” At scale across millions of interactions, the 14x speed improvement translates to lower infrastructure costs and smoother user experiences simultaneously.

Limited Rollout and Capacity Constraints

Ultrafast is currently available in preview form, and OpenAI is restricting access to a small subset of customers. The company stated that it would expand availability as “capacity grows,” a phrasing that acknowledges the infrastructure constraints underlying the feature. Cerebras cannot immediately manufacture chips at scales equivalent to OpenAI’s current global deployment footprint, and the partnership presumably operates under capacity limitations that require gradual expansion.

This staged rollout strategy is standard practice for OpenAI when introducing new capabilities, but it also reflects a genuine technical bottleneck. The company cannot immediately offer 750-token-per-second processing to its entire user base through existing infrastructure. The preview phase serves both as a testing ground and as a way to manage demand against available supply.

Market Implications and Developer Expectations

The announcement of Ultrafast, despite current preview limitations, sends a clear signal to the market about the direction of AI infrastructure investment. Specialized hardware partnerships appear to be the path forward for achieving significant performance gains beyond what general-purpose computing can deliver. This has implications for data center operators, cloud providers, and enterprises evaluating their own AI infrastructure strategies.

For developers, the feature also reframes what is achievable. The existence of a 750-token-per-second configuration of GPT-5.6 Sol means that applications previously considered impractical due to latency—real-time multi-turn conversational systems, live personalization engines, synchronous decision-support tools—become viable target architectures. The feature may eventually reshape what developers consider possible in real-time AI applications.

OpenAI’s announcement arrives as the entire AI industry continues to grapple with scaling. The company has invested heavily in custom silicon partnerships, including arrangements with manufacturers beyond Cerebras, precisely to ensure that inference capacity keeps pace with demand. Ultrafast represents one concrete outcome of those efforts, though it is likely one of several performance-tier options that will eventually come to market.

Frequently Asked Questions

What is Ultrafast and how fast does it work?

Ultrafast is a new mode for GPT-5.6 Sol that OpenAI announced on Thursday. It delivers up to 750 output tokens per second, representing 14x the speed of standard processing for the same model.

Why is Ultrafast limited to preview access right now?

Ultrafast is powered by a partnership with Cerebras, a chipmaker specializing in AI processors. Current capacity limitations mean OpenAI is rolling it out to a small customer group first, with plans to expand access as capacity grows.

What kinds of applications is Ultrafast designed for?

OpenAI identified incident response, customer service and support, financial market analysis, and e-commerce as primary use cases where real-time processing speed creates immediate business value.

Written by
Priya Deshmukh

Priya Deshmukh covers the technology and startup ecosystem — venture capital rounds, founder profiles, and the business models behind the fastest-growing tech companies.