NVIDIA has published its first MLPerf Inference v6.1 results for the Vera Rubin NVL72, reporting leading performance for the system. The announcement frames the result not just as a benchmark win but as evidence of how hardware and software choices shape the cost side of AI inference.

According to NVIDIA, higher system performance produces more tokens per unit of time, which can translate into higher revenue for operators. Efficient scaling, meanwhile, keeps throughput growing as infrastructure expands. The company names continuous software optimization as the third lever, suggesting the gains come from co-design rather than silicon alone.

Because this article draws on a single source, there is no independent or contrasting perspective to compare. The claims reflect NVIDIA's own description of the MLPerf submission.