OpenAI has previewed a new API service tier, Ultrafast, built for GPT-5.6 Sol. The company says the tier runs the model at up to 14 times the speed and can produce as many as 750 output tokens per second. Cerebras hardware underpins the service.
The announcement highlights inference speed as the key differentiator, which matters for applications where latency is a bottleneck. The source does not provide independent performance data or comparisons to other services, so these figures reflect OpenAI's own description.
With only one source available, there are no conflicting reports to weigh. The speed and hardware claims are presented as the company's preview statement rather than an independently verified benchmark.