GPT-5.6 Sol Achieves 14× Speed Boost to 750 Tokens/s with OpenAI’s Ultrafast Mode
OpenAI and Cerebras unveiled an Ultrafast mode for GPT-5.6 Sol that delivers up to 750 tokens per second—14 times faster than the standard baseline—while preserving quality, outpacing rivals such as Claude Fable 5 and Opus 4.8, and reshaping high‑speed AI workflows.
On August 14, OpenAI together with AI‑chip maker Cerebras previewed the Ultrafast mode for its flagship model GPT‑5.6 Sol, offering a maximum output speed of 750 tokens/s, which is about 14 × the 53 tokens/s inference baseline of the standard mode and does not degrade quality. In direct comparisons, the accelerated GPT‑5.6 Sol is 11 × faster than Claude Fable 5 and 5 × faster than Opus 4.8 in Fast mode.
Ultrafast mode is initially released through a limited preview of the OpenAI API.
In the Humanity’s Last Exam (HLE) benchmark—a 2 500‑question test that typically requires PhDs in chemistry, economics, or literature—GPT‑5.6 Sol in Ultrafast mode completed all questions in 11 h 11 m, whereas Claude Fable 5 needed 78 h 27 m, making the former roughly seven times faster while achieving comparable accuracy.
On the GDP‑Val benchmark, which measures economic‑value knowledge‑work, the Ultrafast configuration delivered a 5.6× end‑to‑end speed improvement without quality loss, illustrating how faster inference can accelerate high‑value knowledge tasks.
The speed gains open new workflow possibilities, including:
Event response and reliability: AI can analyze logs, recent code changes, and engineer reports in real time to pinpoint causes and help draft fixes while incidents are still unfolding.
Financial research and security: Rapid analysis of market signals, trade evaluation, and detection of suspicious activity.
Customer support and voice: Multi‑step problem solving without breaking conversational flow.
Commerce: Instant product Q&A, inventory checks, personalized recommendations, and checkout assistance to prevent cart abandonment.
Real‑time research and experimentation: Transforming overnight experiments into interactive work sessions, enabling multiple iterations within a single day.
OpenAI engineers tested Ultrafast mode on incident‑response scenarios: when an alert fires, the model quickly ingests logs, traces information, aggregates dialogue, proposes next checks, and assists in drafting or validating remediation steps, shortening the latency from observation to action while keeping human judgment in the loop.
Researchers also use Ultrafast to swiftly search knowledge bases, query data, and aggregate information from diverse tools, compressing a workflow that previously spanned night‑long runs into a same‑day iterative process.
The breakthrough stems from Cerebras’s wafer‑scale engine (WSE‑3), which integrates 40 trillion transistors, 125 PFLOPS of AI compute, and up to 44 GB of on‑chip SRAM. By keeping model parameters resident in high‑bandwidth SRAM, the architecture eliminates the memory‑bandwidth bottleneck that hampers traditional GPU clusters during large‑model autoregressive decoding.
Because the full parameter set of GPT‑5.6 Sol exceeds a single wafer’s capacity, Cerebras employs a “pipelined‑across‑wafers” strategy that distributes each network layer across multiple wafers, allowing tokens to flow seamlessly between them while each layer’s weights stay in local SRAM.
For many large‑model users, the 14× speed increase means tasks that previously required switching to smaller models (e.g., Luna or Terra) can now run on the flagship model, compressing multi‑hour agent pipelines into minutes.
Faster inference is expected to change how AI is deployed, enabling agents to sit on critical paths of workflows. The AI community is already looking forward to Ultrafast modes for other models such as Luna and Terra, though it remains uncertain whether Cerebras’s hardware can sustain those workloads.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
