OpenAI Tests Ultrafast Mode for GPT-5.6 Sol
OpenAI has begun testing a new Ultrafast mode for its GPT-5.6 Sol model, leveraging Cerebras hardware to deliver speeds up to 14 times faster for latency-critical enterprise applications.

OpenAI has introduced a limited preview of a new Ultrafast execution mode for its GPT-5.6 Sol model through the OpenAI API. Powered by hardware from chipmaker Cerebras, this new mode is capable of running up to 14 times faster than the model's Standard mode. According to OpenAI, the system can generate up to 750 output tokens per second, representing a massive leap in processing velocity for large language models.
Historically, developers and enterprise practitioners had to make a difficult trade-off between model intelligence and latency. Achieving near-instantaneous responses typically required downgrading to smaller, less capable models. The Ultrafast mode aims to eliminate this compromise by pairing high-tier cognitive performance with rapid-fire execution. OpenAI is targeting this capability at tasks where every millisecond counts, such as analyzing system logs and code during active infrastructure outages, or resolving customer issues in real time.
The speed increase also opens up new possibilities for real-time data analysis, rapid hypothesis testing, and instantaneous fraud or suspicious activity detection. Beyond external API customers, OpenAI is actively utilizing Ultrafast mode internally to accelerate its own engineering and research workflows. The company reports that the rapid feedback loop helps its internal teams test ideas and make development decisions much faster than before.
Currently, the Ultrafast mode is restricted to a select group of developers and enterprise customers during its initial testing phase. OpenAI plans to expand access to a broader audience in the future.
This is our own summary of reporting by Mindstream AI



