Test report DSG-8756 · Rev E · tested October 10, 2026

AI Datacenter InfrastructureDevice under test

XCENA MX1 Packs 3,072 RISC-V Cores into a CXL Memory Card

XCENA's MX1, built with Samsung, hosts up to 2 TB of DDR5, 3,072 RISC-V cores, FP16/FP32 vector engines and SSD-as-memory caching on a PCIe 6.0 / CXL 3.2 add-in card consuming 90 W.

Read
3 min
Words
654
Node
28nm
Operator
Priya Raman

Spec summary

  1. 3,072 RISC-V cores at 1.1 GHz on Samsung 4 nm, 40 W chip and 90 W board
  2. Up to 2 TB DDR5 over PCIe 6.0 / CXL 3.2 x8 with 128 GB/s host bandwidth
  3. ~3 TFLOPS FP16/FP32 dot product throughput via subsystem-level vector engines
  4. 20+ TFLOPS advantage over a hypothetical 8.8 TFLOPS GPU bottlenecked on a 1 byte/FLOP CXL link
  5. SSDs exposed as CXL memory with 1,024-entry map cache and 16 GB default capacity

The 3,072 RISC-V cores in XCENA's MX1 — fabricated on Samsung's 4 nm process and clocked at 1.1 GHz — consume 40 W on-chip and produce about 3 TFLOPS of FP16/FP32 dot product throughput.

Jointly developed with Samsung, the MX1 ("MX" for Memory Xcelerator) ships as a PCIe 6.0 / CXL 3.2 x8 add-in card. The card targets a precise market pain: ML models consume ever more memory, and conventional CXL expanders leave DRAM bandwidth on the table that the host cannot harvest.

What does the MX1 actually host?

Up to 2 TB of DDR5 sits in the card's DIMM slots. The upstream CXL link runs 128 GB/s, 64 GB/s each direction. Eight downstream PCIe 6 lanes tether to SSDs, which the device exposes as CXL memory; the attached DDR5 then caches SSD content. XCENA markets the feature as "Infinite Memory."

Compute topology

  • 3,072 in-order RISC-V cores, 1.1 GHz
  • 32 cores per cluster, sharing a 128 KB L2 data cache and a data TLB
  • 4 clusters make one subsystem, the smallest job allocation unit
  • 24 subsystems — 24 independent jobs at once
  • 2 Arm Cortex A53 cores handle housekeeping
  • An in-house NoC ties subsystems to L3 and the DDR5 pool

Each RISC-V core draws under 13 mW. The board with four DIMMs installed consumes 90 W.

How does the cache hierarchy work?

XCENA borrowed a GPU-style layout to suppress translation overhead:

  • 4 KB virtually addressed L1D per core; no translation on L1D hit
  • 128 KB cluster-shared VIPT L2 data cache
  • 8 KB instruction cache shared by 4 cores; 128 KB cluster-L2 instruction cache
  • TLB: 1,024 entries for 64 KB pages, 8 entries for 1 GB pages

Instruction fetches use physical addresses only, so wild jumps to data regions cannot occur. Jobs are isolated at the 128-core subsystem boundary.

The vector engine

RISC-V's extensibility lets XCENA bolt a subsystem-level Vector Processing Engine (VPE) onto each cluster. Each core feeds a per-core VPE command queue with custom instructions. FP32 and FP16 dot products reach roughly 3 TFLOPS chip-wide. At 1.1 GHz, each VPE sustains 128 FLOPS per cycle.

The VPE skips integer math entirely. The 3,072 cores absorb those operations directly, delivering about 3 TOPS at one op per cycle. Built-in VPU functions return explicit error codes; code must check for overflow and invalid accesses manually.

Why near-memory compute here?

The design echoes Intel's Xeon Phi: many weak cores throughput highly parallel work while staying frugal on power. XCENA quantifies the bandwidth bargain against a comparable GPU:

  • An 8.8 TFLOPS-class GPU (GeForce GTX 1080 era) collapses to 64 GFLOPS when forced to load one byte per FLOP over a CXL link
  • The MX1 sustains 200+ GFLOPS in the same scenario, even with cache misses

Compute beside the DDR5 avoids both the bandwidth cap and the power penalty of round-tripping across CXL.

SSDs as memory

XCENA pairs upstream and downstream PCIe bandwidth at roughly 128 GB/s, so an SSD array can saturate the host link without DRAM help. The cache runs on 64 KB pages tracked by a 1,024-entry map cache. Misses trigger page faults that onboard firmware services by fetching from SSD.

Default cache capacity is 16 GB, adjustable in 16 MB steps. Users can also pin a "pinned prefix" of addresses to DRAM contiguously. One documented setup pins 115.5 GB of 231 GB attached DRAM for a shared document workload. MX1 also supports prefetch and SSD RAID to overlap IO with compute.

CXL integration

As a Type 3 (CXL.mem) device, the MX1 can emit back-invalidations to host caches; onboard compute results become visible without forcing memory regions uncacheable. XCENA's software mirrors host page tables, letting host and device share pointers in an OpenCL-SVM style.

via xcena.com (Original)

Filed under

  • cxl
  • risc-v
  • near-memory-computing
  • ddr5
  • samsung
Share this article:

More from Priya Raman

Priya Raman

Show full bio

Correspondent covering business strategy at Die Signal.

243 articles

Same lot · LOT-C1C6

« Previous article