Technology

OpenAI's New Ultrafast Mode Makes Its AI Generate Text 14 Times Faster

Martin HollowayPublished 16h ago4 min readBased on 7 sources
Reading level
OpenAI's New Ultrafast Mode Makes Its AI Generate Text 14 Times Faster
source:openai.com

OpenAI has released a preview of Ultrafast, a new way to use its GPT-5.6 Sol AI model that runs up to 14 times faster than normal, producing up to 750 words per second. The preview is powered by OpenAI's partnership with chipmaker Cerebras and is currently available to a small group of customers, with access set to expand as capacity grows. OpenAI

For anyone building software that uses AI, the speed number is what matters most. AI text models generate their output in small pieces called "tokens," which are roughly equivalent to words or parts of words. At 750 tokens per second, Ultrafast is fast enough that the time between a user asking a question and getting an answer becomes so short that a person would barely notice any waiting at all. For comparison, GPT-5.6 Sol's existing Fast mode delivers up to 2.5x faster speeds than Standard processing at twice the price, with no change in intelligence, according to OpenAI's earlier pricing and performance disclosures. Ultrafast jumps from 2.5x to 14x over the standard speed.

OpenAI is positioning Ultrafast for use in business settings including incident response, customer service and support, financial market analysis, and e-commerce, according to TechCrunch. These are situations where speed is not just about user comfort but about whether the system can actually do its job. An automated assistant that can generate repair steps for a technical outage in under a second changes what companies can realistically automate. A financial-analysis tool that produces results at 750 tokens per second can deliver insights almost as events unfold.

The model itself, GPT-5.6 Sol, is described by OpenAI as a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with the company's most advanced safety measures. It achieves top results across coding and knowledge work, per OpenAI's model announcement. An August 5 update to GPT-5.6 Sol in ChatGPT delivered more focused answers and adapted its level of detail to the question being asked, according to OpenAI.

The competitive landscape adds useful context. Anthropic, which makes the rival Claude AI, offers a "fast mode," but it does not deliver the speed that Ultrafast offers, TechCrunch reports. The 14x speed boost, if it holds up in everyday use, puts OpenAI in a different performance tier from what Anthropic currently offers.

The Cerebras partnership is what makes this possible. Cerebras builds specialized computer chips designed specifically to run AI models quickly, rather than the more general-purpose chips that power most AI systems today. The fact that OpenAI is using a third-party chipmaker for this premium service rather than running it on its own systems is worth noting. It suggests that the demands of Ultrafast are best met by hardware purpose-built for speed.

OpenAI's August 13 announcement was accompanied by a companion publication, "The builder's guide to GPT-5.6," under its Applied AI category, suggesting the company is investing in helping developers use the technology alongside the raw performance claim. The same day brought news of Dali Rajic's appointment as Chief Revenue Officer. Earlier in the week, OpenAI confirmed it is testing ads in ChatGPT (August 11), made Daybreak models available on AWS (August 11), and published two security posts on August 10 expanding its Daybreak cyber defense initiatives and putting frontier cyber models in more trusted hands. An August 12 company post titled "How enterprises put AI to work" rounded out a week of business-facing messaging.

Looking at what this means in practice, the speed jump from Fast to Ultrafast is not a small step forward. It crosses a line where the AI model's output speed stops being the thing that slows everything down. At that point, the limits move to other parts of the system: how quickly relevant information can be gathered and sent to the model, and how fast the software receiving the AI's output can act on it. For teams building on GPT-5.6 Sol, the question is no longer whether the model is fast enough for their needs but whether the rest of their software can keep up.

The limited preview scope is the practical caveat. OpenAI has not disclosed pricing for Ultrafast, nor a timeline for when it will be widely available, nor how the speed holds up under heavy, sustained use versus short bursts. What is known is that the tier exists, it runs at 14x standard speed on Cerebras hardware, and it is being tested with a small set of customers. Anyone considering building around it should treat those unknowns as open questions.