Test report DSG-1507 · Rev B · tested October 10, 2026

Processors & AcceleratorsDevice under test

NextSilicon Ships Maverick-2 Dataflow Engine, Unveils Arbel RISC-V Core

After $303M and eight years of development, NextSilicon ships its 54B-transistor Maverick-2 dataflow engine, targeting HPC with HBM3E memory, a self-tuning compiler, and a new Arbel RISC-V core.

Read
4 min
Words
711
Node
65nm
Operator
Amara Osei

Spec summary

  1. NextSilicon ships Maverick-2 after 8 years and $303M in funding; Maverick-1 launched in 2022 with Sandia National Laboratories.
  2. Maverick-2 die: 54B transistors at TSMC 5nm, 224 compute blocks, 32 RISC-V E-cores, ~1.5 GHz clock, HBM3E memory.
  3. Single-die TDP 400W; dual-die OAM module 750W (both raised from 300W/600W disclosed last year).
  4. Published benchmarks: 32.6 GUPS at 460W (22× CPU, 6× GPU), 5.2 TB/sec STREAM, 600 GFLOPS HPCG matching leading GPU at half power, 10× PageRank.
  5. Arbel RISC-V core: 10-wide issue, 6 integer ALUs, 4×128-bit vector FPUs, 64KB L1i/d, 1MB L2, claimed parity with Intel LionCove and AMD Zen 5.

After eight years of development and $303 million in seed and venture funding, NextSilicon has begun shipping its 64-bit Maverick-2 dataflow engine, a 54-billion-transistor die manufactured on TSMC's 5nm process. The company also disclosed Arbel, a homegrown RISC-V processor it labels a test chip.

The launch follows the 2022 Maverick-1 proof of concept, which NextSilicon developed with Sandia National Laboratories. Sandia is widely expected to take delivery of the first production Maverick-2 system.

How does the dataflow engine differ from a CPU?

NextSilicon's architecture, branded the Intelligent Computing Architecture (ICA), replaces the Von Neumann fetch-decode-execute loop with a grid of compute blocks that an application's intermediate representation maps itself onto. Ilan Tayari, co-founder and VP of architecture, framed the change in stark silicon-area terms.

"Today's high-end processors have become complicated and chunky, both physically and practically," Tayari said at the launch. "They dedicate 98 percent of their silicon to overhead, traffic management, data shuffling — not actual computation."

The Maverick-2 die carries 224 compute blocks in a seven-by-eight grid, with 32 RISC-V E-cores flanking the left and right edges. NextSilicon does not publish the ALU count per block, but its slide deck shows a 14-by-14 internal grid, suggesting 196 ALUs per block and roughly 44,000 ALUs across the die. Threads run at 1.5 GHz. CEO Elad Raz said the runtime can load and delete "mill cores" — graph-shaped thread bundles — in nanoseconds as workloads shift.

What are the package-level specs?

Two Maverick-2 packages ship today:

  • Single-die: 400W TDP, raised from 300W disclosed last year
  • Dual-die OAM: 750W TDP, raised from 600W

Both pair with HBM3E memory. The dual-die module delivers 5.2 TB/sec of measured bandwidth, or 83.9 percent of peak, NextSilicon said.

How does it benchmark against CPUs and GPUs?

NextSilicon published four benchmark figures at the launch:

  • GUPS: 32.6 GUPS at 460W — claimed 22× CPU, 6× GPU
  • STREAM: 5.2 TB/sec, 1.86× perf-per-watt versus a GPU
  • HPCG: 600 gigaflops at 600W — "matching leading GPU performance" at half the power
  • PageRank: 10× improvement over "leading GPUs"

Peak FP64 on a single Maverick-2 sits near 11 TF. That compares with 33.5 TF on Nvidia H100 vector cores and 67 TF on H100 tensor cores. Raz has argued publicly that sustained throughput, not peak numbers, drives purchasing decisions.

What does the compiler do?

The Maverick compiler takes C, C++, or Fortran source, lifts its IR onto the compute blocks, and continuously re-tunes the resulting dataflow at runtime. Hot code paths — typically 80 percent of runtime — move to the ALU blocks; the remainder runs on the embedded RISC-V E-cores or the host X86 CPU. No porting to CUDA-X or ROCm is required.

"NextSilicon's dataflow architecture allows us to achieve significantly lower overhead compared to traditional CPUs and GPUs," Tayari said. "We pivot the silicon allocation ratio. We dedicate the majority of the resources to actual computation rather than control overhead."

What is Arbel?

Arbel is NextSilicon's first RISC-V CPU, built around a homegrown core with the following characteristics:

  • 10-wide issue decoder
  • 6 integer ALUs
  • 4 vector FPUs at 128-bit width
  • 16 scalar instructions in parallel
  • 64 KB L1 instruction cache, 64 KB L1 data cache
  • 1 MB L2 cache, 2 MB cache per core

NextSilicon claims the Arbel core can "stand toe-to-toe" with Intel's LionCove Xeon core and AMD's Zen 5 Epyc core. The company has not published core count or clock speed for the Arbel chip.

What does the launch change in HPC procurement?

Maverick-2 targets HPC buyers first, with AI inference and training as a secondary workload. The 750W OAM module competes with Nvidia H100 and H200 SXM accelerators in the same socket class but delivers a dataflow runtime that ports CPU code automatically and self-tunes for the duration of a run.

Sandia remains the flagship reference customer. Whether Arbel graduates from test chip to production silicon will shape whether NextSilicon can ship a fully homegrown host-accelerator pair without depending on third-party X86 or Arm CPUs.

via en.wikipedia.org (Original)

Filed under

  • nextsilicon
  • maverick-2
  • risc-v
  • dataflow-architecture
  • hpc-accelerators
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

254 articles

Same lot · LOT-C1C6

« Previous article