Test report DSG-8299 · Rev D · tested September 30, 2026

AI Datacenter InfrastructureDevice under test

CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production

CoreWeave brings NVIDIA Vera Rubin NVL72 to production with Cognition as first customer, reporting 4.8x token throughput over GB200, plus Vera CPU and the new Forge platform.

Read
4 min
Words
846
Node
5nm
Operator
Marcus Bennett

Spec summary

  1. Vera Rubin NVL72 delivered up to 4.8x higher token throughput than GB200 NVL72 in Cognition's SWE-2 inference tests.
  2. CoreWeave's Vera CPU rack holds 128 CPUs and 11,264 cores, supporting over 11,000 concurrent isolated agent environments with 3x faster sandbox startup.
  3. CoreWeave Forge combines Weights & Biases, OpenPipe and marimo; serverless RL trains 1.4x faster at 40% lower cost than self-managed setups.

CoreWeave announced production availability of NVIDIA Vera Rubin NVL72 systems paired with Spectrum-X 102.4T Ethernet networking at its Fully Connected event in San Francisco this week, making it one of the first cloud providers to deliver the platform to customers. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.

The announcements extend a co-engineering relationship between the two companies that spans nearly a decade. CoreWeave will also offer NVIDIA Vera, billed as the first CPU built for AI agents, and launched CoreWeave Forge, a connected environment for training, evaluating and improving models and agents on NVIDIA accelerated computing.

"NVIDIA accelerated computing delivers value across generations," said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA. "CoreWeave's NVIDIA V100 GPUs are still running customer workloads nearly a decade after Volta launched, even as CoreWeave brings Vera Rubin NVL72 into production. That's the strength of the NVIDIA platform: infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload."

Cognition Benchmarks: 4.8x Token Throughput

Cognition runs training, reinforcement learning and production inference for Devin on CoreWeave, having scaled to thousands of GPUs on the platform in nine months. Earlier this month, CoreWeave received its first Vera Rubin NVL72 production racks. Cognition then benchmarked Vera Rubin's inference performance against a GB200 NVL72 baseline using a real-world software engineering workload, sampling tasks from FrontierCode and deploying AI agents to solve them.

In early tests, Vera Rubin NVL72 delivered up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72. For Devin, that translates into faster real-time code generation and more responsive multistep reasoning.

"Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship," said Silas Alberti of Cognition's founding team. "Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec."

CoreWeave stood up a production Vera Rubin cluster for Cognition in days, citing codesign and collaboration with NVIDIA across the stack, from infrastructure to tokens served. Customers can operate the capacity through CoreWeave Kubernetes Service, SUNK, CoreWeave Mission Control, CoreWeave Sandboxes and CoreWeave Inference.

Vera CPU: 11,264 Cores Per Rack, 3x Faster Sandbox Startup

Agentic AI strains infrastructure in two directions: serving agents demands low-latency compute at scale, while post-training requires thousands of isolated environments running simultaneously.

CoreWeave's Vera deployment packs 128 CPUs and 11,264 cores into a single rack — enough for more than 11,000 concurrent environments at one core each. With CoreWeave Sandboxes, these environments are hardware-isolated and run alongside the training jobs they support, with Spectrum-X Ethernet switches and BlueField-4 DPUs handling agent communication at low latency.

In testing, CoreWeave achieved more than 3x faster agent sandbox startup times on Vera CPUs. On Terminal-Bench, CoreWeave recorded a 1.7x performance gain across all passing tasks.

CoreWeave Forge Closes the Loop

CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project into one environment for continuous model and agent improvement, remaining open across models, frameworks and clouds. New and expanded capabilities include:

  • CoreWeave ARIA, now generally available, analyzes runs, proposes experiments, recommends code changes stored in GitHub, and surfaces what drove a change to propose the next experiments.
  • CoreWeave Agent Lens, a new service, improves failure detection by 20% and fixes issues at half the cost, turning tens of millions of production agent traces into insights.
  • CoreWeave Sandboxes, now generally available, run agents, tool calls, RL and evaluations in isolated CPU or GPU environments — serverless or on existing training infrastructure.
  • Serverless post-training: serverless supervised fine-tuning and serverless RL use customers' production signals with no training cluster required. Serverless RL trains 1.4x faster at 40% lower cost than a self-managed setup.

NVIDIA Dynamo, the open source inference framework for AI factories, powers CoreWeave's managed inference service and RL Rollouts, now in private preview. RL Rollouts loads new checkpoints into a live deployment while running, so reinforcement learning continues without redeploys.

Canva, Capital One and MasterClass are among the first companies building on Forge. NVIDIA Nemotron open models give Forge teams a path to customizing and deploying reasoning and multimodal models for agentic workflows.

Enterprise Traction

Ennoble Care, a home-based care provider serving about 50,000 high-need Medicare patients across 15 states, selected CoreWeave to run clinical AI inference. It will use reserved NVIDIA RTX PRO 6000 GPU capacity on CoreWeave Kubernetes Service to scale AI agents for clinical documentation, decision support and back-office automation.

CoreWeave has delivered record MLPerf results in every round of training and inference, holds the Platinum ranking in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0 — the only cloud provider to do so — and serves nine of the 10 leading AI labs.

via nvidia.com (Original)

Filed under

  • nvidia
  • coreweave
  • vera-rubin
  • ai-infrastructure
  • inference
Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

News editor covering marketplaces and e-commerce at Die Signal.

58 articles

Same lot · LOT-C1D0

« Previous articleNext article »