NVIDIA said on Sept. 16 that its first Vera Rubin NVL72 preview submission in MLPerf Inference v6.1 reached up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x higher throughput on DeepSeek-R1. The company presented the results as evidence of faster inference economics, but the public record is narrower than that framing.

MLCommons said the new v6.1 release adds benchmarks for agentic and end-to-end inference and describes MLPerf Inference as architecture-neutral, representative and reproducible. It also says the published results are intended to give procurement teams empirical data for AI system decisions. That makes MLPerf relevant here — but only within the limits of the submitted workloads and disclosure terms.

Those limits matter. NVIDIA’s Vera Rubin material covers only two benchmarks in preview status: DeepSeek-R1 and Qwen3-VL. The company says the Qwen3-VL result spans offline, server and interactive scenarios, and that the DeepSeek-R1 result used TensorRT-LLM. MLCommons’ own results analysis adds another qualification: the largest v6.1 gains in VLM and DeepSeek-R1 came from a new preview category system powered by Vera Rubin. That is a useful indicator of where the platform stands, but it is not the same as proving a broad production advantage across unrelated workloads.

The submission does establish a specific point. Vera Rubin NVL72 entered MLPerf Inference with strong scores on two of the suite’s most demanding newer tests, and it outpaced GB300 NVL72 on both in the submitted configurations. NVIDIA also says the results reflect full-stack codesign and ongoing software optimization, while noting that some later performance gains were not yet verified by MLCommons. Those are vendor claims, not independent proof of real-world deployment outcomes.

What buyers still cannot infer from this material is just as important: pricing, supply, power draw in their own environment, integration effort, or total cost per token outside the benchmark setup. The benchmark also does not answer whether a customer’s specific model mix, traffic pattern or latency target would see the same gap.

The cleanest conclusion is limited but concrete: Vera Rubin NVL72 has a strong MLPerf debut on two preview benchmarks, and MLCommons’ updated suite makes those tests more relevant to agentic and multi-step inference than older single-shot measures. The next procurement-relevant signal is whether Vera Rubin appears in additional verified MLPerf submissions beyond these two preview workloads, because that would widen the evidence base from a narrow debut to a broader benchmark record.