Test report DSG-8756 · Rev E · tested October 10, 2026
AI Datacenter InfrastructureDevice under test
XCENA MX1 Packs 3,072 RISC-V Cores into a CXL Memory Card
XCENA's MX1, built with Samsung, hosts up to 2 TB of DDR5, 3,072 RISC-V cores, FP16/FP32 vector engines and SSD-as-memory caching on a PCIe 6.0 / CXL 3.2 add-in card consuming 90 W.
- Read
- 3 min
- Words
- 654
- Node
- 28nm
- Operator
- Priya Raman
Spec summary
- 3,072 RISC-V cores at 1.1 GHz on Samsung 4 nm, 40 W chip and 90 W board
- Up to 2 TB DDR5 over PCIe 6.0 / CXL 3.2 x8 with 128 GB/s host bandwidth
- ~3 TFLOPS FP16/FP32 dot product throughput via subsystem-level vector engines
- 20+ TFLOPS advantage over a hypothetical 8.8 TFLOPS GPU bottlenecked on a 1 byte/FLOP CXL link
- SSDs exposed as CXL memory with 1,024-entry map cache and 16 GB default capacity
The 3,072 RISC-V cores in XCENA's MX1 — fabricated on Samsung's 4 nm process and clocked at 1.1 GHz — consume 40 W on-chip and produce about 3 TFLOPS of FP16/FP32 dot product throughput.
Jointly developed with Samsung, the MX1 ("MX" for Memory Xcelerator) ships as a PCIe 6.0 / CXL 3.2 x8 add-in card. The card targets a precise market pain: ML models consume ever more memory, and conventional CXL expanders leave DRAM bandwidth on the table that the host cannot harvest.
What does the MX1 actually host?
Up to 2 TB of DDR5 sits in the card's DIMM slots. The upstream CXL link runs 128 GB/s, 64 GB/s each direction. Eight downstream PCIe 6 lanes tether to SSDs, which the device exposes as CXL memory; the attached DDR5 then caches SSD content. XCENA markets the feature as "Infinite Memory."
Compute topology
- 3,072 in-order RISC-V cores, 1.1 GHz
- 32 cores per cluster, sharing a 128 KB L2 data cache and a data TLB
- 4 clusters make one subsystem, the smallest job allocation unit
- 24 subsystems — 24 independent jobs at once
- 2 Arm Cortex A53 cores handle housekeeping
- An in-house NoC ties subsystems to L3 and the DDR5 pool
Each RISC-V core draws under 13 mW. The board with four DIMMs installed consumes 90 W.
How does the cache hierarchy work?
XCENA borrowed a GPU-style layout to suppress translation overhead:
- 4 KB virtually addressed L1D per core; no translation on L1D hit
- 128 KB cluster-shared VIPT L2 data cache
- 8 KB instruction cache shared by 4 cores; 128 KB cluster-L2 instruction cache
- TLB: 1,024 entries for 64 KB pages, 8 entries for 1 GB pages
Instruction fetches use physical addresses only, so wild jumps to data regions cannot occur. Jobs are isolated at the 128-core subsystem boundary.
The vector engine
RISC-V's extensibility lets XCENA bolt a subsystem-level Vector Processing Engine (VPE) onto each cluster. Each core feeds a per-core VPE command queue with custom instructions. FP32 and FP16 dot products reach roughly 3 TFLOPS chip-wide. At 1.1 GHz, each VPE sustains 128 FLOPS per cycle.
The VPE skips integer math entirely. The 3,072 cores absorb those operations directly, delivering about 3 TOPS at one op per cycle. Built-in VPU functions return explicit error codes; code must check for overflow and invalid accesses manually.
Why near-memory compute here?
The design echoes Intel's Xeon Phi: many weak cores throughput highly parallel work while staying frugal on power. XCENA quantifies the bandwidth bargain against a comparable GPU:
- An 8.8 TFLOPS-class GPU (GeForce GTX 1080 era) collapses to 64 GFLOPS when forced to load one byte per FLOP over a CXL link
- The MX1 sustains 200+ GFLOPS in the same scenario, even with cache misses
Compute beside the DDR5 avoids both the bandwidth cap and the power penalty of round-tripping across CXL.
SSDs as memory
XCENA pairs upstream and downstream PCIe bandwidth at roughly 128 GB/s, so an SSD array can saturate the host link without DRAM help. The cache runs on 64 KB pages tracked by a 1,024-entry map cache. Misses trigger page faults that onboard firmware services by fetching from SSD.
Default cache capacity is 16 GB, adjustable in 16 MB steps. Users can also pin a "pinned prefix" of addresses to DRAM contiguously. One documented setup pins 115.5 GB of 231 GB attached DRAM for a shared document workload. MX1 also supports prefetch and SSD RAID to overlap IO with compute.
CXL integration
As a Type 3 (CXL.mem) device, the MX1 can emit back-invalidations to host caches; onboard compute results become visible without forcing memory regions uncacheable. XCENA's software mirrors host page tables, letting host and device share pointers in an OpenCL-SVM style.
via xcena.com (Original)
More from Priya Raman
Same lot · LOT-C1C6
- DSG-138910nmIntel's Crescent Island Puts 480 GB of LPDDR5X on a Datacenter GPU
- DSG-27037nmCXMT's Fifth-Generation DRAM Platform Enters Mass Production
- DSG-93523nmXsight Labs E1L Scales 200Gbps DPU Down to 25W Envelope
- DSG-90773nmCXMT Starts Mass Production of G5 DRAM as HP and Dell Qualify Parts
- DSG-386920nmd-Matrix Pairs Raptor XPU with Nvidia NVL144 for 144-Node Inference Rack