OpenAI has launched GPT-6 Astra Ultrafast, a faster inference mode built on NVIDIA Blackwell GPUs and available now through the OpenAI API and to eligible ChatGPT Work and Codex users. According to NVIDIA, Ultrafast offers up to 8x faster token generation than Astra Standard mode, a jump that matters most when responses are repeated across multi-step workflows such as agents writing code, calling tools, and checking results.
The speedup is aimed at making interactive applications feel more responsive and shortening the time between tool calls in coding agents. NVIDIA credits OpenAI's inference optimizations, which tap into Blackwell's architecture, for the acceleration. OpenAI's inference lead, Philippe Tillet, said in the announcement that NVIDIA's investment in tooling and documentation helped OpenAI make its models particularly good at programming NVIDIA GPUs, and that Astra can turn that knowledge into high-performance kernels.
Performance gains are expected to continue after deployment. OpenAI is using its own models to help refine the inference software running on NVIDIA GPUs, taking advantage of the platform's programmability to test and implement improvements. Uday Ruddarraju, OpenAI's CTO of compute, said this ongoing work helped deliver the acceleration behind Astra Ultrafast and can make deployed infrastructure more productive over time. NVIDIA also highlights that a programmable platform lets teams reuse infrastructure across training, inference, and reinforcement learning, improving utilization as demand shifts.