Test report DSG-4056 · Rev E · tested October 10, 2026
AI Datacenter InfrastructureDevice under test
NVIDIA Vera Rubin NVL72 Delivers Up to 30x Throughput Per Megawatt
Vera Rubin NVL72 hits up to 30x throughput per megawatt on SemiAnalysis AgentX; Lambda gains 23% performance per watt with DSX MaxLPS. AI Infra Summit, Santa Clara.
- Read
- 4 min
- Words
- 831
- Node
- 28nm
- Operator
- Marcus Bennett
Spec summary
- Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on SemiAnalysis AgentX (DeepSeek V4 Pro), with up to 45x lower cost per million tokens.
- Lambda ran 19 nodes in the power budget of 16, raising token throughput 24% (about 4M to 5M tokens/second) and performance per watt 23% with DSX MaxLPS.
- In its first MLPerf Inference preview submission, Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72; a 288-GPU GB300 submission reached 99% scaling efficiency.
- Groq 3 LPX with Vera Rubin delivers up to 35x token throughput per megawatt vs GB200 NVL72 for 2T+ parameter models at long context; 2,529 output tokens/second per user on Qwen 3.8 27B at 100K context.
- AI Infra Summit drew more than 8,000 attendees, up from 3,500 last year; Emerald AI and Silicon Valley Power demonstrated automated demand-response across hundreds of grid signals.
NVIDIA's Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on real-world agentic workloads, according to SemiAnalysis AgentX results published Tuesday at the AI Infra Summit in Santa Clara. On the DeepSeek V4 Pro model, the figure reaches up to 30x per megawatt, with the AgentX results showing up to 45x lower cost per million tokens.
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, presented the results before more than 8,000 attendees, up from 3,500 last year. "Infrastructure that's fungible, that's reliable, that's going to last 10 years and really becomes an asset for computing the world's computing problems in the world's industries — they can build that with Vera Rubin, with DSX and all the innovations that we have here," Buck said.
The metric for AI infrastructure is shifting from peak performance to validated agentic tokens per megawatt. NVIDIA claims DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization.
What did Lambda measure with DSX MaxLPS?
AI cloud provider Lambda provided the first validation of NVIDIA DSX MaxLPS on Blackwell servers. The numbers:
- 19 nodes run within the power budget typically allocated to 16 full-power nodes
- Cluster-wide token throughput up 24%, from about 4 million to 5 million tokens per second
- Performance per watt up 23%
- Up to 40% more GPU capacity within the same megawatt budget on Vera Rubin NVL72, in the right deployment environments
DSX MaxLPS continuously monitors power consumption across GPUs and racks, shifting available power where demand is highest and reclaiming capacity static provisioning leaves unused.
What did MLPerf Inference v6.1 show?
In its first MLPerf Inference preview submission, Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72. A 288-GPU submission spanning four GB300 NVL72 racks achieved 99% scaling efficiency, with throughput growing nearly linearly from a single-rack baseline.
Software alone drove up to 1.6x higher performance in v6.1 submissions versus v6.0, with additional gains achieved after the benchmark submission period. MLCommons peer-reviews every result before publication.
How does Groq 3 LPX fit in?
NVIDIA pairs Vera Rubin with Groq 3 LPX deterministic ultralow-latency inference for agentic AI, where agents chain reasoning steps and tool calls. The combined platform delivers up to 35x higher token throughput per megawatt than GB200 NVL72 for 2-trillion-plus-parameter models at long context.
On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX hit 2,529 output tokens per second per user. At the factory level, the platform offers up to 40% more GPUs within the same site-power envelope and up to 35% higher token throughput without new power lines.
The SemiAnalysis results reflect the shape of agentic workloads: a single session can accumulate hundreds of thousands of input tokens — roughly 15x the volume of a simple chat request — as agents reason, call tools and spawn sub-agents. The performance stems from the NVL72 scale-up domain with sixth-generation NVLink, NVFP4 precision on fifth-generation Tensor Cores, and an inference stack spanning TensorRT LLM and Dynamo. Vera Rubin is in full production, NVIDIA said.
Can AI factories act as flexible grid resources?
Emerald AI and NVIDIA demonstrated a commercial flexible-load program with Silicon Valley Power. The system responded to hundreds of demand signals from the utility while protecting AI workload performance.
Emerald AI plans to use NVIDIA DSX Flex in Conductor, its grid-responsive power management software. DSX Flex receives load-shedding requests, demand-response events and pricing signals, and automatically throttles back power on low-priority jobs while critical workloads keep running.
What are partners and startups reporting?
New collaborations announced at the summit:
- Amazon's Annapurna Labs is working with NVIDIA on NVHBM custom high-bandwidth memory
- d-Matrix integrates with NVLink Fusion, combining NVIDIA Vera CPUs with d-Matrix Raptor XPUs for ultralow-latency inference at scale
- Pinterest uses the Blackwell platform and Dynamo inference software for conversational AI in visual discovery
Startups benchmarked the NVIDIA Vera CPU:
- Perplexity: 1.9x faster sandbox starts for its SPACE secure sandbox platform
- DeepInfra: wins on all metrics, including 2.2x faster orchestration step latency
- Redpanda: 5.5x lower latencies and 73% higher throughput than other CPUs
- Starburst: 3x faster query throughput
- Kinetica: 2.7x faster analytical query performance
- ClickHouse: fastest machine measured so far on ClickBench — a "strong signal of what's ahead for CPU performance for data-intensive workloads," the company said
- Daytona cited "serious gains for agentic execution"; Prime Intellect reported sustained high bandwidth and consistent memory latency under parallel load
NVIDIA also detailed NVLink 6's multilayer resiliency architecture for factories scaling to hundreds of thousands of GPUs. Custom forward error correction, physical layer retry and universal physical layer recovery maintain a lossless fabric, while credit-based flow control, dynamic routing and link rebalancing contain faults locally and prevent cascading stalls.
via ai-infra-summit.com (Original)
More from Marcus Bennett
Show full bio
News editor covering marketplaces and e-commerce at Die Signal.
276 articles
Same lot · LOT-C1C6
- DSG-460814nmFoxconn Puts NVIDIA Vera Rubin AI Datacenter Cost at $47 Billion Per Gigawatt
- DSG-82995nmCoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production
- DSG-369465nmNVIDIA Ships DGX Spark 64GB at $4,999 for Local AI
- DSG-55495nmAlibaba Targets 20 GW Data Center Capacity by 2032 Backed by V900
- DSG-831028nmAMD Buys Taalas to Add Model-Specific Silicon to Instinct Roadmap