Test report DSG-2010 · Rev A · tested October 10, 2026

Processors & AcceleratorsDevice under test

Fujitsu Details 144-Core Monaka Arm CPU at Hot Chips 2026

Fujitsu revealed its 144-core Monaka Arm CPU at Hot Chips 2026: 2nm compute dies stacked on 5nm I/O and SRAM base dies, with DDR5 memory, 2.9 GHz clocks, and 6 ALU pipes.

Read
3 min
Words
600
Node
5nm
Operator
Marcus Bennett

Spec summary

  1. Monaka packs 144 Arm cores across four 2nm compute dies stacked on a 5nm I/O and SRAM base die.
  2. Clock speed is 2.9 GHz versus A64FX's 2 GHz, but trails current AMD, Intel, and Arm Neoverse server cores.
  3. Core carries six ALU pipes, a 64 KB instruction cache, and a 'three-level' TAGE branch predictor.
  4. Vector pipes shrank from A64FX's two 512-bit units to two 256-bit units; A64FX maxed at 32 GB HBM2.
  5. Per-core area measures 1.47 mm² versus Neoverse N2's 1 to 1.3 mm², with roughly double the per-core vector throughput.

Fujitsu's Monaka CPU delivers 144 Arm cores split across four 2nm compute dies stacked on a 5nm I/O and SRAM base die. The company presented the chip at Hot Chips 2026 as the successor to A64FX, the 52-core, 2 GHz design that powered the Fugaku supercomputer.

What changes from A64FX?

The Monaka core pairs a "three-level" TAGE branch predictor—Fujitsu's own terminology for cascading tables indexed with progressively longer global history—with a 64 KB instruction cache feeding a decoupled fetch pipeline. A64FX relied on a 2048-entry Branch Weight Table implementing a perceptron-style predictor. AMD last used perceptrons as the primary prediction method after Zen 1.

Monaka's backend carries six ALU pipes versus A64FX's four. Six pipes imply at least 6-wide issue, since narrower issue could not sustain those ports. The vector units shrank from A64FX's two 512-bit pipes to two 256-bit pipes. The narrower width trades raw vector throughput for power and area, since most Arm code targets NEON at 128 bits while SVE benefits from simpler, narrower datapaths.

Three power-saving mechanisms target the vector units:

  • A register cache in front of the FP register file holds recently read values, similar to GPU register reuse caches.
  • Predicate pattern detection skips reads when predicates contain common patterns such as all-1s.
  • Renamer zero-tracking avoids useless reads or writes to the upper 128 bits when NEON code leaves them zero.

How does the chiplet design work?

The hub-and-spoke package mirrors AMD's Zen 5 VCache arrangement. Each 2nm compute die stacks on a 5nm base die holding last-level cache and LDOs. Fujitsu places SRAM on the older node because SRAM scaling has slowed. Analog power delivery does not benefit from process shrinks, so LDOs stay on the base die. The hotter compute die sits on top, closer to cooling.

Fujitsu built a custom low-voltage SRAM macro to keep the base die below standard Vmin. Slides from ICS 2024 suggest an assist circuit that raises voltage during access.

Which NUMA configurations does Monaka offer?

A64FX exposed a fixed four-cluster NUMA topology tied to HBM2. Monaka, now on DDR5, offers three modes:

  • 8 nodes: each compute die splits into two nodes of 18 cores plus half the L3, the lowest-latency option.
  • 4 nodes: one node per compute die, tied to the nearest memory controllers, comparable to AMD EPYC's NPS4 mode.
  • 1 node: the whole chip appears as one domain, letting non-NUMA-aware software scale across all 144 cores.

DDR5 swaps A64FX's 32 GB HBM2 ceiling for removable modules with much higher capacity.

What security and RAS features does Monaka add?

Monaka ships with cloud- and enterprise-grade features beyond HPC norms:

  • Hardware root of trust (TPM-style) for boot integrity
  • Memory encryption for cold-boot resistance
  • Pointer authentication to block return-oriented programming
  • CLRBHB and related mitigations against Spectre variants

For reliability, Fujitsu brought mainframe-class RAS from its supercomputer lineage. Cache ECC plus hardware instruction retry let the core flush at commit, reload known-good architectural state, and re-execute the failing instruction in single-step mode before returning to out-of-order operation.

What remains undisclosed?

Fujitsu did not specify L3 capacity, reorder buffer entries, or register file sizes. The 2.9 GHz clock trails current AMD, Intel, and Arm Neoverse server cores. That gap will cap single-thread throughput. At 1.47 mm² per core, Monaka sits slightly above Neoverse N2 at 1 to 1.3 mm², yet offers roughly twice the per-core vector throughput.

via substackcdn.com (Original)

Filed under

  • fujitsu
  • monaka
  • arm
  • hot-chips
  • chiplets
Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

News editor covering marketplaces and e-commerce at Die Signal.

278 articles

Same lot · LOT-C1C6

« Previous articleNext article »