Product Launch2026-08-14OpenAI Blog

OpenAI previews Ultrafast mode for GPT-5.6 Sol

OpenAI has announced a preview of Ultrafast, a new API service tier that dramatically accelerates the performance of GPT-5.6 Sol. Powered by Cerebras, this mode delivers up to 750 output tokens per second, making it up to 14 times faster than standard configurations. The service is specifically designed for enterprise users who require high-speed, low-latency responses from their AI systems. The introduction of Ultrafast addresses a critical need in the enterprise market. Many applications, such as real-time customer support, financial trading, and interactive content generation, demand immediate responses. Delays of even a few seconds can impact user experience and operational efficiency. With Ultrafast, OpenAI is providing a solution that meets these stringent requirements. Cerebras, the hardware partner behind this initiative, is known for its specialized AI processors that excel at high-throughput tasks. By integrating Cerebras technology into the OpenAI API, the two companies have created a service that pushes the boundaries of what is possible in terms of speed. This collaboration is a testament to the importance of hardware innovation in advancing AI capabilities. For enterprises, the benefits of Ultrafast are clear. Faster response times enable more dynamic interactions, allowing businesses to deploy AI in scenarios where it was previously impractical. Additionally, the high token throughput means that larger, more complex models can be used in real-time applications without sacrificing performance. While Ultrafast is currently in preview, OpenAI plans to expand availability based on demand. Early adopters will have the opportunity to test the service and provide feedback, helping to shape its final features. As AI continues to evolve, services like Ultrafast will play a crucial role in making advanced models accessible and practical for enterprise use.

Related news