< Back to all clusters
[TECHNOLOGY] · United States · 27 sources

started · updated

OpenAI launches Ultrafast mode for GPT-5. 6 Sol powered by Cerebras

OpenAI has introduced Ultrafast, a new API service tier designed to run its GPT-5. 6 Sol model at significantly higher speeds. Powered by CerebrasWafer-Scale Engine technology, the mode can generate up to 750 output tokens per second, which is up to 14 times faster than standard processing.

This advancement aims to eliminate the traditional trade-off between model intelligence and response speed. Previously, developers requiring real-time performance often had to opt for smaller, less capable models. Ultrafast allows access to the full reasoning capabilities of GPT-5. 6 Sol with ultra-low latency. The service is currently in a limited preview for a select group of API customers, including companies in sectors such as financial research, coding, customer support, and e-commerce.

In performance benchmarks, Cerebras reported that the Ultrafast mode completed the Humanity’s Last Exam—a 2,500-question PhD-level test—in approximately 11 hours and 11 minutes. This was notably faster than comparable performance from Anthropic’s Claude models. OpenAI is also utilizing the tier internally for time-sensitive tasks like incident response and log analysis.

Entities

Andrew Feldman · Anthropic · Cerebras · GPT-5. 6 Sol · OpenAI · Sachin Katti

Claims

What the coverage asserts, and how well corroborated each claim is across sources.

Sources

4 days ago
GPT-5. 6 Slashes Agent Costs [www.startuphub.ai]
5 days ago