OpenAI has released initial performance figures for Jalapeño, its custom inference chip. The company says the chip delivers industry-leading speed and efficiency for running modern AI models, with higher throughput and lower latency than current alternatives.
A key focus of the design is power efficiency. By tailoring the silicon to inference workloads, OpenAI aims to reduce the energy cost of serving models while maintaining fast response times. The results represent the first public look at the chip's capabilities.
As with any vendor-reported benchmark, these figures have not been independently verified. Still, the announcement signals OpenAI's push to control more of its hardware stack, particularly for the compute-intensive inference stage that dominates production AI costs.