OpenAI Launches 'Ultrafast' Mode for GPT-5.6 Sol, Boosting Speed by 14x
On August 13, 2026, OpenAI launched "Ultrafast," a new mode for its GPT-5.6 Sol model, enabling it to run up to 14 times faster by delivering up to 750 output tokens per second. This significant speed increase, powered by a partnership with chipmaker Cerebras and its Wafer-Scale Engine architecture, is initially available in a limited preview to select customers through the OpenAI API, opening up real-time use cases for frontier-level AI. For broader context, explore our AI News.
Ultrafast Mode: A Leap in Processing Speed
The core of the Ultrafast mode is its ability to deliver up to 750 output tokens per second, a substantial increase over previous speeds. This acceleration is attributed to OpenAI's partnership with Cerebras, leveraging their Wafer-Scale Engine (WSE) architecture. The WSE is designed to provide high-throughput inference, presenting a notable alternative to traditional Nvidia GPUs in specific high-performance scenarios.
Cerebras Wafer-Scale Engine Powers Performance
Cerebras's Wafer-Scale Engine is central to the performance gains observed with Ultrafast mode. This specialized hardware integrates 44 GB of SRAM onto a single wafer-sized chip, optimizing it for demanding AI workloads. The collaboration highlights a strategic move to enhance the underlying infrastructure supporting advanced AI models, focusing on efficiency and speed for complex computational tasks.
Real-World Implications and Benchmarks
The introduction of Ultrafast mode is expected to open new possibilities for real-time AI applications. OpenAI suggests potential use cases in areas such as incident response, customer support, financial analysis, and cybersecurity, where immediate processing of information is critical. The ability to perform frontier-level reasoning at interactive speeds could transform how businesses integrate AI into their operational workflows.
Cerebras provided benchmark data illustrating the performance of Ultrafast mode. In a test involving 2,500 questions from "Humanity's Last Exam," Ultrafast completed the task in approximately 11 hours. This contrasts with over 78 hours required by Claude Fable 5. Furthermore, Cerebras reports that Ultrafast runs 11 times faster than Fable 5 and 5 times faster than Opus 4.8 in its fast mode, all while maintaining comparable quality.
Availability and Future Outlook
Ultrafast mode is currently in a limited preview phase, accessible to a select group of customers through the OpenAI API. This controlled rollout allows OpenAI to gather feedback and refine the service before a broader release. The partnership with Cerebras and the successful implementation of the Wafer-Scale Engine architecture suggest a growing trend towards specialized hardware solutions for accelerating AI inference, potentially diversifying the market beyond dominant GPU providers.
Conclusion
OpenAI's launch of Ultrafast mode for GPT-5.6 Sol marks a significant advancement in AI processing speed, driven by its collaboration with Cerebras. By achieving up to 14 times faster performance and delivering 750 output tokens per second, this new mode aims to unlock real-time applications for complex AI tasks. As the limited preview progresses, the industry will observe how this technology reshapes the landscape of high-throughput AI inference and its practical applications.
Sources
Recommended AI tools
Perplexity
Search & Discovery
Clear answers from reliable sources, powered by AI.
Cursor
Code Assistance
The AI code editor that understands your entire codebase
Google Cloud Vertex AI
Data Analytics
Gemini, Vertex AI, and AI infrastructure—everything you need to build and scale enterprise AI on Google Cloud.
Adobe Firefly
Image Generation
Create your way with Adobe Firefly—AI for every creative vision.
Google AI Studio
Productivity & Collaboration
The fastest way to build AI-first applications with Google Gemini.
Hugging Face
Scientific Research
Democratizing good machine learning, one commit at a time.
Was this article helpful?
Found outdated info or have suggestions? Send us a note.