Test report DSG-4687 · Rev F · tested October 10, 2026

Processors & AcceleratorsDevice under test

AMD's Instinct MI455X Pairs 432 GB HBM4 with 72-GPU Helios Rack

AMD's Instinct MI455X packs 432 GB of HBM4 and 40.26 PFLOP at MXFP4 per GPU on CDNA5, while the Helios rack scales to 72 GPUs, 260 TB/s of UALoE bandwidth, and 2.9 ExaFLOPS per cabinet.

Read
5 min
Words
942
Node
65nm
Operator
Amara Osei

Spec summary

  1. Single MI455X delivers up to 40.26 PFLOP at OCP MXFP4 and 315 TFLOP at FP32 matrix/vector
  2. Each GPU pairs 432 GB of HBM4 across 12 stacks with 23.3 TB/s of memory bandwidth at 2.4 GHz
  3. Helios rack holds 72 GPUs for 2.9 ExaFLOPS at MXFP4, 22.6 PFLOP at FP32, and 31 TB of shared HBM4
  4. Aggregate L2 bandwidth reaches 54 TB/s across two 96 MB Fabric Cache Dies linked at 14 TB/s bidirectional
  5. CDNA5 retires the 15-year GCN lineage in favor of an RDNA-derived compute unit
AMD’s Instinct MI455X: Aiming for the Sun
Fig. AAMD’s Instinct MI455X: Aiming for the Sun — AI-generated

AMD's Instinct MI455X delivers 40.26 PFLOP at OCP MXFP4 per GPU and pairs that figure with 432 GB of HBM4 delivering 23.3 TB/s of memory bandwidth, the company said at its Advancing AI event. The chip replaces the Instinct MI355X at the top of AMD's AI stack and ships as the first GPU designed from the ground up for rack-scale AI deployments.

Built on the new CDNA5 architecture and packaged in TSMC's CoWoS-L, the MI455X carries 256 Work Group Processors distributed across 8 Accelerator Complex Dies (XCDs). AMD set the maximum engine clock at 2.4 GHz, with peak figures ranging from 315 TFLOP for FP32 matrix or vector and vector FP16, up to 40.26 PFLOP at OCP MXFP4.

What does the MI455X change at the compute unit?

CDNA5 reorganizes the Compute Unit. AMD now counts Work Group Processors (WGPs) instead of CUs, the same terminology the company adopted for RDNA4. Each WGP contains four dual-issue Wave32 SIMD32 units alongside four scalar units, replacing CDNA4's four single-issue Wave64 SIMD16 units. The result: each WGP executes 256 packed FP32 operations per cycle, or 512 FLOPS using FMA.

The vector register file scales accordingly. Any wave can now address up to 1,024 VGPRs, four times the prior CDNA and RDNA limit. Each SIMD still holds 128 KB of vector registers (half RDNA4's 192 KB), but the move to Wave32 effectively doubles the registers visible to a single thread, to 1,024. A single wavefront can occupy the entire SIMD.

Matrix throughput climbed as well. Each matrix unit handles 8,192 FP4 operations per cycle. With four matrix units per WGP, a single WGP delivers 65,536 matrix operations per cycle. AMD acknowledges that "only FP4 and FP8 are practically faster in the new architecture"; the rest of the gain comes from the SIMD widening from 16 to 32 lanes.

L1 and LDS capacity doubled. The L1 Data Cache now sits at 64 KB; the LDS at 320 KB. Each cache can deliver two 256-byte transfers per clock, doubling prior bandwidth.

How does CDNA5 restructure the cache hierarchy?

Each XCD holds 32 active WGPs (34 physical, two fused off), split into two Shader Engines of 16 WGPs each. AMD replaced the per-SA GL1 buffer from RDNA4 with a per-SE Broadcast Arbitrator that doubles as a write-combine buffer and provides up to 4x bandwidth amplification by replicating data to multiple WGPs.

The memory subsystem introduces multicast loads. In AMD's example, a GEMM that computes C = A × B normally forces each wavefront to load its tile of A into its own LDS. With multicast loads, the package fetches A once from L2 and the Broadcast Arbitrator replicates it into every relevant WGP, while each wavefront still loads its own B tile.

Below the WGPs sit two Fabric Cache Dies (FCDs) holding 96 MB of 1 MB SRAM blocks each, for 192 MB of Global L2 across the package. Each FCD drives up to 27 TB/s of L2 bandwidth, an aggregate of 54 TB/s. CDNA5 removes cross-die L2 access, a feature CDNA3 and CDNA4 offered. AMD framed the change as improving atomic behavior, eliminating the kernel-flush boundary MI300-series XCDs required, and increasing L2 data reuse. The two L2 domains stay isolated, but each can hold data for any global address, so a WGP's local L2 can cache data tied to its sibling FCD without crossing the die-to-die link.

That die-to-die link between FCDs carries roughly 14 TB/s of bidirectional bandwidth. Each FCD also attaches to six stacks of HBM4 on a 2,048-bit interface. Each stack holds 36 GB, totaling 432 GB at 23.3 TB/s. Pin speed sits at approximately 7.6 GT/s.

For external I/O, an IO die on each FCD connects the package to the host CPU through a dedicated 16-lane AMD Infinity Fabric link delivering 256 GB/s bidirectional, replacing PCIe. AMD describes the link as coherent between the GPU and the host.

How does the MI455X scale out to 72 GPUs per rack?

The MI455X carries 36 400-Gbit/s UALoE ports, each implementing two 200G Ethernet lanes, for 3.6 TB/s of peak bidirectional scale-up bandwidth per GPU. AMD pairs the silicon with Helios, its rack-scale platform built on OCP's Open Rack Wide form factor (1.2 m wide, 1.3 m deep).

A Helios cabinet holds 18 Compute Trays arranged in two groups. Each tray hosts four MI455X GPUs and one 96-core EPYC 9006 SP7 host CPU with 16 × 64 GB DIMMs (1 TB) and five E1.S SSD slots. Six Switch Trays provide 12 UALoE switches, each with 432 200 Gb/s links (21.6 TB/s per switch), reached through three UALoE links per GPU. Aggregate scale-up bandwidth reaches 260 TB/s bidirectional per rack.

Scale-out comes from Pensando NIC boards configurable with four or six NICs, delivering up to 43 TB/s of backend bandwidth. The full 72-GPU rack supports up to 2.9 ExaFLOPS at MXFP4 or 22.6 PFLOP at FP32, with 31 TB of shared HBM4 at 1.7 PB/s of memory bandwidth.

AMD framed CDNA5 as the end of the GCN microarchitecture lineage that began with Tahiti, ran through Fiji, Vega, and the first four CDNA generations, and now hands off to a RDNA-derived design. With Helios, the company extends that transition from silicon to rack, pairing EPYC CPUs, GFX12 IP, and Pensando NICs into a single validated deployment.

via substackcdn.com (Original)

Filed under

  • amd-instinct-mi455x
  • hbm4
  • cdna5
  • helios
  • tsmc-cowos-l
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

254 articles

Same lot · LOT-C1C6

« Previous articleNext article »