Test report DSG-3869 · Rev D · tested October 10, 2026
AI Datacenter InfrastructureDevice under test
d-Matrix Pairs Raptor XPU with Nvidia NVL144 for 144-Node Inference Rack
d-Matrix will place its Raptor XPU inside Nvidia's NVL144 MGX rack with 144 accelerators and 7.2 PB/sec aggregate bandwidth per chassis, the company said on Thursday.
- Read
- 3 min
- Words
- 591
- Node
- 20nm
- Operator
- Amara Osei
Spec summary
- Raptor slot: 144 XPUs per Nvidia NVL144 MGX rack, 7.2 PB/sec aggregate memory bandwidth per chassis
- Per-card spec: 4nm TSMC compute die on custom DRAM at 36-micron pitch, 100 TB/sec bandwidth, 2.3 TB 3D-stacked DRAM
- Schedule: Raptor tape-out by end of 2026, release in Q4 2027; Corsair shipped in June claiming 10x Nvidia GPU inferencing speed
- Funding: d-Matrix has raised more than $500M including from Microsoft M12 since its 2019 founding
- Market: Nvidia CFO Colette Kress put AI factory revenue at $18B per gigawatt on 2022 Grace-Hopper and $40B per gigawatt on forthcoming Vera-Rubin systems

Seven years after founding with an exclusive focus on AI inferencing, d-Matrix will place its forthcoming Raptor XPU inside Nvidia's NVL144 MGX rack, the company said this week. The Thursday announcement packs 144 memory-centric accelerators per chassis and brings 7.2 petabytes per second of aggregate memory bandwidth to inference workloads.
d-Matrix has raised more than $500 million, including capital from Microsoft's M12 venture arm, before shipping its first product. The startup, founded in 2019, released Corsair in June with executives claiming the memory-centric platform can run inferencing workloads 10 times faster than Nvidia GPUs alone.
What does the Raptor platform deliver?
Detailed at Hot Chips 2026, Raptor integrates a 4-nanometer compute die manufactured by Taiwan Semiconductor Manufacturing Co atop a custom DRAM die at a 36-micron pitch. Each card produces 100 TB/sec of memory bandwidth. d-Matrix expects tape-out by the end of 2026 and release in the fourth quarter of 2027.
Lightning, a follow-on platform with a multi-high DRAM stack, remains on the roadmap and will fall under the Nvidia partnership.
How does the Nvidia partnership work?
Raptor will plug into Nvidia's MGX reference architecture and ship as part of the AI factory offering. The companies will share NVLink switch trays and compute trays with Vera-Rubin systems. Each Raptor card sits alongside Nvidia Vera CPUs and Bluefield DPUs, with ConnectX and Spectrum-X handling scale-out.
Per-rack technical specifications:
- 144 Raptor XPUs
- 2.3 TB of 3D-stack DRAM at 100 TB/sec per card
- 7.2 PB/sec aggregate memory bandwidth
- Standard NVL144 liquid-cooling plumbing
Operators can pair a Raptor rack with adjacent Vera-Rubin GPU racks rather than substitute for them.
What performance did d-Matrix show at Hot Chips?
d-Matrix benchmarked two models on Raptor during the conference. Z.ai's GLM 5.2 produced about 3,000 tokens per second per user. Moonshot AI's Kimi K3 reached 1,000 tokens per second per user. CEO Sid Seth said d-Matrix can scale these results across eight racks.
How is Nvidia positioning the deal?
Nvidia framed the partnership as another instantiation of its open-platform approach. Jesse Clayton, principal product marketing manager for Nvidia's Data Center GPU business, called the AI factory stack "completely fungible" and "vertically integrated, but horizontally open." NVLink Fusion ports already admit third-party silicon such as Groq's 3 LPX accelerator racks.
The deal also plugs d-Matrix into a fast-growing market. CFO Colette Kress, presenting Nvidia's Q2 2027 earnings, said the AI factory revenue opportunity has grown from about $18 billion per gigawatt on 2022's Grace-Hopper rackscale systems to $40 billion per gigawatt on the forthcoming Vera-Rubin systems.
What sits behind d-Matrix's resource thesis?
Seth, who co-founded the company in 2019, framed the company's economics at this week's briefing.
"Our entire approach is predicated on doing more with the capital people deploy in our compute," Seth said. "We are able to run really, really fast compute. We do more inference with a little amount of time, and we have made a very energy-efficient solution with the memory-centric computing. That allows us to do more with less of these resources – money, time, and energy. We can hopefully, over time, alleviate the need to build out more datacenters."
d-Matrix acquired GigaIO's data-center business in April and shipped a SquadRack reference design built with Arista Networks, Broadcom, and Supermicro.
via hotchips.org (Original)
More from Amara Osei
Same lot · LOT-C1C6
- DSG-617865nmNvidia Plans Investment in Inference Chip Startup d-Matrix
- DSG-117645nmNvidia Plans to Invest in AI Chip Rival d-Matrix
- DSG-891845nmNvidia to Invest in Chip Rival d-Matrix as Its Challengers Choose Partnership
- DSG-42317nmNvidia and SK Group Detail $500B-Plus AI Data Center, Memory Push
- DSG-369465nmNVIDIA Ships DGX Spark 64GB at $4,999 for Local AI