OpenAI previews ultrafast mode for GPT-5.6 Sol at 750 tokens per second

OpenAI has introduced a limited preview of Ultrafast mode for GPT-5.6 Sol, saying the new tier can deliver up to 14 times faster output than the model’s standard version and reach 750 output tokens per second versus a roughly 53 tokens per second baseline. Announced on August 13, the offering is initially available to a select group of customers through the OpenAI API, with broader access planned as computing capacity expands. The service builds on OpenAI’s existing Fast mode, which speeds GPT-5.6 Sol by up to 2.5 times, but Ultrafast is positioned as a separate tier enabled by a partnership with Cerebras and its wafer-scale chip architecture. OpenAI and Cerebras said GPT-5.6 Sol in Ultrafast mode completed Humanity’s Last Exam, a 2,500-question benchmark, in 11 hours and 11 minutes, compared with 78 hours and 27 minutes for Anthropic’s Claude Fable 5. OpenAI also said Ultrafast delivered a 5.6x end-to-end speed improvement on GDP-Val without quality degradation. The companies said the model runs about 11 times faster than Claude Fable 5 and roughly 5 times faster than Opus 4.8 in Fast mode. Early users include Jane Street, Basis, Rogo and Podium, with OpenAI highlighting applications in incident response, financial research, security operations, voice AI, e-commerce, and research workflows where lower latency can materially change how the model is used.

The information on this website is generated using AI and we cannot guarantee its accuracy. Please use it as reference information only.