Test report DSG-9865 · Rev F · tested October 10, 2026
AI Datacenter InfrastructureDevice under test
CoreWeave opens Vera Rubin NVL72 to customers, adds Ennoble Care
CoreWeave now offers Nvidia's Vera Rubin NVL72 rack-scale system on its cloud platform. Coding-agent developer Cognition reports a 4.8x total token throughput gain on SWE-2 inference workloads since early September.
- Read
- 3 min
- Words
- 681
- Node
- 7nm
- Operator
- Elena Vasquez
Spec summary
- Cognition recorded a 4.8x total token throughput gain on SWE-2 inference workloads, with deployment beginning in early September.
- The NVL72 rack combines 36 Vera CPUs and 72 Rubin GPUs and cuts installation time from two hours to five minutes via cable-free modular trays.
- Nvidia claims Rubin delivers 5x inference and 3.5x training performance versus the prior Blackwell generation.
- CoreWeave will offer Vera as a standalone bare-metal CPU at 128 CPUs and 11,264 cores per rack, with customer testing starting in the coming weeks.
- CoreWeave also signed Ennoble Care to CoreWeave Kubernetes Service on RTX Pro 6000 Blackwell Server Edition nodes and launched CoreWeave Forge in general availability at its San Francisco conference.
CoreWeave now offers Nvidia's Vera Rubin NVL72 rack-scale system on its cloud platform, with coding-agent developer Cognition running production workloads on the hardware since early September.
The two companies confirmed the deployment at CoreWeave's Fully Connected conference in San Francisco. Cognition recorded a 4.8x increase in total token throughput on its SWE-2 inference workloads against the prior generation.
What does the NVL72 system contain?
The NVL72 rack combines 36 Vera CPUs and 72 Rubin GPUs. Each Rubin platform module pairs six chip types:
- NVLink 6 Switch
- ConnectX-9 SuperNIC
- BlueField-4 DPU
- Spectrum-6 Ethernet Switch
- Vera CPU
- Rubin GPU
Vera succeeds Nvidia's Grace CPU. Rubin succeeds Blackwell GPUs. Nvidia claims Rubin delivers 5x inference and 3.5x training performance versus Blackwell.
How does the rack differ from the Blackwell era?
The Rubin NVL72 is 100% liquid-cooled and uses cable-free modular tray designs. Nvidia states the new design reduces rack installation from two hours to five minutes.
Who is the first customer?
Cognition, developer of the SWE agent family, began using CoreWeave's Vera Rubin systems in early September.
"Bringing up Nvidia Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations," said Chen Goldberg, executive vice president of product and engineering at CoreWeave.
"With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls, and thousands of concurrent tasks put pressure on the entire platform."
Goldberg added: "Our job is to make compute, networking, and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves."
Silas Alberti, SVP research and founding team at Cognition, confirmed the throughput numbers. "We have seen up to a 4.8x increase in total token throughput for our SWE-2 inference workloads on the new generation," Alberti said.
Will Vera CPUs be sold on their own?
CoreWeave plans to offer Vera as a standalone bare-metal CPU service. At rack scale, CoreWeave will deliver 128 Vera CPUs and 11,264 cores per rack, paired with BlueField-4 DPUs and Spectrum-X Ethernet switches.
"General-purpose infrastructure bottlenecks agentic AI; Vera is the first CPU explicitly designed to accelerate it," Goldberg said. "Our platform natively enables Vera with products like CoreWeave Sandboxes out of the box."
Goldberg added: "Teams can instantly spin up thousands of isolated environments, removing operational friction and accelerating the entire AI loop on day one."
Customer testing starts in the coming weeks, according to Corey Sanders, SVP of product at CoreWeave. A small group has had early access.
What else shipped at Fully Connected?
CoreWeave signed home-care provider Ennoble Care to CoreWeave Kubernetes Service. Ennoble Care deploys Nvidia RTX Pro 6000 Blackwell Server Edition nodes for summarization, documentation, and clinical decision support.
"We've built a full-stack, ONC-certified EMR purpose-built for home-based primary care, and we're now developing multiple AI agents on top of it, both to augment clinical delivery and to automate back-office functions," said Jonathan Taylor, CTO of Ennoble Care.
Taylor added: "CoreWeave gives us reserved capacity we can count on, a Kubernetes environment our team can move into quickly, and engineers who answer the phone. That combination lets us extend AI capabilities already in our workflows to tens of thousands of additional patients."
Jon Jones, chief revenue officer at CoreWeave, framed the win in workload terms. "Clinical inference is one of the most demanding places AI can run. The workload is continuous, the latency budget is short, and the compliance requirements are absolute. That is what The Essential Cloud for AI means in practice."
What is CoreWeave Forge?
The company also launched CoreWeave Forge, a development layer on top of its cloud platform. CoreWeave says Forge lets teams build, improve, and evaluate AI models and agents in a single environment. The product reached general availability at the conference.
via Data Center Dynamics (Source)
More from Elena Vasquez
Show full bio
Senior reporter covering industry trends and analytics at Die Signal.
247 articles
Same lot · LOT-C1C6
- DSG-82995nmCoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production
- DSG-831028nmAMD Buys Taalas to Add Model-Specific Silicon to Instinct Roadmap
- DSG-405628nmNVIDIA Vera Rubin NVL72 Delivers Up to 30x Throughput Per Megawatt
- DSG-55275nmCisco and Nvidia target enterprise AI with rack-scale systems
- DSG-671165nmMicrosoft, AMD Unveil Helios Rack for Azure AI Inference