OpenAI Introduces 'Ultrafast' Mode for GPT-6.1 Sol, Delivering Faster Token Generation at Six Times the Standard Price

In the weeks after OpenAI's DevDay event, the company had already released several models in its GPT-6 series. 

GPT-6 Astra served as the flagship, priced at $10 per million input tokens and $50 per million output tokens. A week later came GPT-6 Sol, followed almost immediately by GPT-6.1 Sol as an upgrade within the same mid-tier slot. 

OpenAI described GPT-6.1 Sol as delivering performance close to Astra on agentic coding, computer use, and professional work while charging one-fifth the standard token rates. Its list price settled at $2 per million input tokens, $10 per million output tokens, and $0.10 per million cached input tokens, with a context window of roughly 1.05 million tokens and a knowledge cutoff of April 30, 2026. 

Then, OpenAI announced the 'Ultrafast' mode, saying that it is rolling out for GPT-6.1 Sol in the API, Codex, and ChatGPT Work. 

Ultrafast is not a separate model. 

It is essentially a service tier that routes requests for the same underlying model onto a lower-latency serving path, reducing the time between successive generated tokens. 

OpenAI stated that the tier can produce tokens up to eight times faster than the standard mode for GPT-6.1 Sol, with some reports and documentation referencing speeds around 300 tokens per second in Codex. 

Its documentation notes that the figure measures token generation rather than end-to-end task completion. 

Network overhead, tool calls, and other external waits remain unchanged, so the observed reduction in total time for an agentic workflow is typically smaller than the raw token-speed improvement. 

In other words, the model itself stays the same. Output quality, context length, reasoning effort options, and other capabilities are unchanged. 

What changes is how the request is scheduled and executed so that tokens are generated faster.

The company recommends WebSockets for multi-turn agentic applications to preserve more of the latency gain.

Pricing for the Ultrafast tier is 6x the standard rates, or $12 per million input tokens and $60 per million output tokens for prompts within the ordinary context range.

Reactions on developer forums and social platforms have been mixed. 

Some users noted the value for interactive loops where a person is waiting between steps, such as live debugging or steering an agent in real time. Others pointed out that the six-times price and restricted plan access make the tier impractical for routine or high-volume work, and that included usage on even the Pro plan can be exhausted quickly when the higher rate applies. 

Comments also emphasized that the practical benefit depends on measuring full task time rather than token rate alone, since many coding and agent workflows spend substantial time outside model generation. 

It's worth noting that earlier public version of Ultrafast (announced in August 2026 for GPT-5.6 Sol), was powered by Cerebras wafer-scale engines. 

Those systems keep model weights in large amounts of on-chip SRAM rather than repeatedly fetching them from off-chip high-bandwidth memory, which removes a major bottleneck in conventional GPU inference and enables very high token rates. 

That version was described as reaching up to 750 output tokens per second.⁠

Published