OpenAI has launched a limited preview of "Ultrafast" mode for its GPT-5.6 Sol model, promising speeds up to 14 times faster than its standard processing. This accelerated performance is achieved through a partnership with Cerebras, which provides its wafer-scale engine architecture to power the new tier. The Ultrafast mode can generate up to 750 output tokens per second, making it suitable for time-sensitive and mission-critical applications.

The development addresses a long-standing trade-off in AI development between model intelligence and response speed. Previously, users often had to choose between the advanced capabilities of larger models and the quicker responses of smaller, less capable ones. OpenAI's Ultrafast mode aims to eliminate this compromise, offering the full intelligence of GPT-5.6 Sol at speeds that can integrate directly into real-time workflows. This is particularly relevant for enterprise use cases where latency can significantly impact the value of AI-driven solutions.

OpenAI is positioning Ultrafast mode for applications such as real-time voice assistants, financial research, customer support, and security breach prevention. The increased speed allows for more natural conversational interactions, faster analysis of financial data, and more immediate responses in critical operational scenarios. For instance, in customer support, a faster model can make automation feel more responsive, while in incident response, speed is directly tied to the effectiveness of the solution. Benchmarks provided by Cerebras indicate that GPT-5.6 Sol on Ultrafast mode performed significantly faster than other leading models on complex tasks. On the "Humanity's Last Exam" benchmark, which consists of PhD-level questions, GPT-5.6 Sol Ultrafast completed the task in approximately 11 hours, compared to over 78 hours for another leading model.

The speed gains are attributed to Cerebras's wafer-scale engine architecture, which packs a large amount of on-chip SRAM. This design aims to overcome the memory bandwidth bottlenecks common in GPU-based systems, where model weights must be repeatedly transferred. By keeping model weights on-chip, Cerebras's hardware minimizes data movement, enabling faster inference. This technological approach is a departure from conventional GPU designs, which can be limited by the speed of data transfer between on-chip memory and off-chip storage.

The Ultrafast mode is currently available in a limited preview to a select group of API customers. OpenAI plans to expand access as its capacity grows. This phased rollout allows the company to gather insights on where the speed improvements create the most value before wider deployment. The partnership between OpenAI and Cerebras represents a significant step in their collaboration, with OpenAI having previously committed $10 billion to Cerebras for low-latency compute.

This launch follows OpenAI's recent efforts to improve AI accessibility and affordability. In July 2026, the company introduced lower prices for its GPT-5.6 Luna and Terra models and faster performance for GPT-5.6 Sol in its API through a "Fast mode." Fast mode offered up to 2.5 times faster speeds than standard processing for GPT-5.6 Sol at double the price. The introduction of Ultrafast mode signifies a further push towards making advanced AI capabilities more readily available for demanding enterprise workloads.