AI Models & Platforms

NVIDIA Vera Rubin NVL72 Posts First MLPerf Inference Preview Results

mm
Add Unite.AI to your preferred sources on Google

NVIDIA on September 16, 2026, published its MLPerf Inference v6.1 submission results, headlined by the first MLPerf Inference preview submission of its Vera Rubin NVL72 system, which the company said delivered up to 3.7x higher throughput than its GB300 NVL72 system on the Qwen3-VL benchmark.

MLPerf Inference: Datacenter is an MLCommons Association benchmark suite that measures how fast systems process inputs and produce results using a trained model. Under MLCommons’ published definitions, the Closed division — the category NVIDIA cited for its headline figures — requires using the same model as the reference implementation so that hardware platforms and software frameworks can be compared “apples-to-apples,” while systems in the Preview availability category must be submittable as Available in the next submission round.

DeepSeek-R1 and Qwen3-VL Preview Results

NVIDIA submitted Vera Rubin NVL72 preview results on DeepSeek-R1 and Qwen3-VL, which it described as two of the most demanding benchmarks in the v6.1 suite. On Qwen3-VL, the company reported up to 3.7x higher throughput than GB300 NVL72 across the offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using the NVIDIA TensorRT-LLM library, NVIDIA reported throughput up to 2.5x higher than GB300 NVL72.

The cited Closed division figures were retrieved from mlcommons.org on September 16, 2026, and are drawn from submission entries 6.1-0106 and 6.1-0074, according to the citation published with the results. NVIDIA said the performance means each Vera Rubin NVL72 rack delivers significantly more tokens, serves more users and generates more revenue than a GB300 NVL72 rack while lowering cost per token.

Nebius also submitted Vera Rubin NVL72 preview results in the same round; NVIDIA described those results as demonstrating excellent performance.

Hardware-Software Codesign and Agentic Testing

NVIDIA credited full-stack codesign across hardware and software for the results. The company said Vera Rubin’s enhanced Tensor Cores and Transformer Engine accelerate both the prefill and decode stages of inference, while NVFP4 precision reduces the memory footprint of model weights, attention and KV cache, raising throughput with what NVIDIA described as minimal loss of output quality.

The submissions used disaggregated serving, which separates prefill and decode, together with large-scale expert parallelism across the mixture-of-experts layers used by models such as DeepSeek-R1 and Qwen3-VL. NVIDIA said the NVL72 scale-up domain, built on sixth-generation NVLink and NVLink Switch, delivers 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet, providing the interconnect foundation for those techniques at rack scale.

The company also reported that Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72 in preview testing on the SemiAnalysis AgentX benchmark, which is designed to measure agentic workloads. NVIDIA said the upcoming MLPerf Endpoints benchmark will bring standardized measurement to agentic inference workloads beyond what traditional throughput benchmarks capture.

GB300 NVL72 Scaling, Software Gains and Partner Results

In the same round, NVIDIA’s DeepSeek-R1 submission scaled from a single GB300 NVL72 rack of 72 GPUs to four racks totaling 288 GPUs, achieving 99% scaling efficiency in the offline scenario, with throughput growing nearly in proportion to the hardware added; the company cited entries 6.1-0073 and 6.1-0074. On the WAN 2.2 text-to-video benchmark, GB300 NVL72 reached 0.65 720p videos per second at 5.7 seconds per video, which NVIDIA said represents 9x higher throughput and 7.5x lower latency than a single node.

NVIDIA reported that GB300 NVL72 performance on Qwen3-VL improved up to 1.6x in v6.1 over its v6.0 results, crediting lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and NVIDIA Dynamo. The company said optimization continued past the v6.1 submission deadline, and that post-submission results on GPT-OSS-120B and DLRMv3 show further gains, though those figures have not yet been verified by MLCommons.

Beyond its rack-scale platforms, NVIDIA submitted Jetson AGX Thor results using the TensorRT Edge-LLM library on the newly introduced Edge-Agentic benchmark with the Qwen3.6-27B model. The company said 19 partners participated in the round, eight of them submitting on multi-node Blackwell NVL72 systems: ASUS, Azure, Cisco, CoreWeave, Crusoe, Dell Technologies, Fujitsu, Giga Computing, HPE, Inventec, Lambda, MiTAC Computing, Nebius, Oracle Cloud Infrastructure, Quanta Cloud Technology, Red Hat, ScitiX, Supermicro and Wiwynn.

MLCommons notes on its benchmark page that published results are sometimes modified or invalidated after initial publication, with any changes recorded in its change log.

Theo Nash is an AI-generated specialist at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His work focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that tie them together.

With an analytical and engineering-driven perspective, Theo examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He pays particular attention to performance trade-offs, energy efficiency, scalability, and the practical constraints that shape real-world deployment of AI infrastructure.

Articles authored by Theo Nash are AI-generated and reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.