NVIDIA's new Vera Rubin NVL72 platform leads in latest MLPerf Inference v6.1 testing, achieving up to 3.7x throughput gains compared to previous generations.
Leading Performance in MLPerf v6.1
NVIDIA recently unveiled the performance capabilities of its new Vera Rubin NVL72 system, establishing a new benchmark in the latest MLPerf Inference v6.1 suite. This marks the first preview submission for the architecture, which is purpose-built to address the demanding economics of AI inference. In rigorous testing, the Vera Rubin NVL72 demonstrated substantial efficiency gains, delivering up to 3.7 times greater throughput compared to the existing GB300 NVL72 systems. This jump in performance is critical for AI infrastructure providers, as increased system throughput directly correlates to higher token generation rates and improved revenue potential per unit of infrastructure. The results highlight the company's commitment to optimizing both hardware and software stacks to support advanced workloads.
Technical Advancements Driving Throughput
The performance improvements are rooted in a full-stack codesign approach that leverages new hardware features and software optimizations. Specifically, the Vera Rubin architecture utilizes enhanced Tensor Cores and the Transformer Engine to accelerate both prefill and decode stages of inference. Furthermore, the introduction of NVFP4 precision effectively reduces the memory footprint required for model weights, attention, and the KV cache. By minimizing this footprint, the system can significantly increase throughput with minimal impact on output quality. The platform also utilizes large-scale expert parallelism, a technique specifically advantageous for Mixture-of-Experts (MoE) models such as DeepSeek-R1 and Qwen3-VL, which were central to the latest benchmark submissions.
Scalability and Interconnect Architecture
Achieving rack-scale performance requires not just individual GPU capability, but also high-speed, low-latency interconnects. The Vera Rubin NVL72 utilizes the sixth-generation NVIDIA NVLink and NVLink Switch, which are designed to deliver 10 times higher packet rates and 3 times lower latency compared to standard off-the-shelf Ethernet solutions. This interconnect foundation is vital for disaggregated serving, where prefill and decode operations are separated to maximize efficiency. Beyond the Vera Rubin debut, testing on the GB300 NVL72 platform demonstrated the resilience of this scaling strategy, achieving 99% scaling efficiency when expanding from a single rack to four racks in DeepSeek-R1 testing scenarios, proving that hardware additions translate almost linearly into performance gains.
Evolving AI Inference Benchmarks
As AI development shifts toward more complex agentic behaviors—where models must reason, plan, and execute tasks across multiple steps—traditional throughput benchmarks are being augmented. The recent MLPerf testing included results that address this shift, such as the SemiAnalysis AgentX benchmark. In these specialized tests, the Vera Rubin NVL72 delivered a 30-fold performance increase over the GB300 NVL72. Looking forward, the introduction of the MLPerf Endpoints benchmark will likely provide a more standardized measurement for agentic inference workloads. This evolution in testing methodologies reflects the broader industry transition from simple token generation to interactive AI agents that require more sophisticated compute infrastructure.
Continuous Software Optimization
Hardware performance is only one component of the equation; software development plays an equally important role in maximizing inference economics. NVIDIA reported that software optimizations within the v6.1 suite alone delivered up to 1.6 times higher performance compared to v6.0 results. These gains were realized through techniques such as additional kernel fusion, the implementation of better kernels, and the use of the NVIDIA Dynamo open-source inference framework alongside vLLM. The company noted that these efforts are continuous, with post-submission optimizations on models like GPT-OSS-120B and DLRMv3 already showing further potential improvements. This ongoing cycle of software refinement ensures that deployed infrastructure continues to extract more value over time.
Implications for Infrastructure Providers
For data center operators and organizations making infrastructure investment decisions, these MLPerf results provide critical data regarding long-term cost and efficiency. By increasing the number of tokens generated per rack, the Vera Rubin NVL72 lowers the cost per token, making high-end AI services more economically viable at scale. The wide ecosystem participation, with 19 partners contributing to these benchmark results, indicates a mature pipeline for deploying these systems globally. As companies balance the need for increased AI compute power with energy constraints, the ability to achieve higher scaling efficiency and throughput per watt becomes the defining factor for the next generation of AI factories and intelligent infrastructure deployments.
No comments:
Post a Comment