Test report DSG-3784 · Rev C · tested October 10, 2026

Processors & AcceleratorsDevice under test

AWS Ships Graviton5: 192 Cores, 2.4x Performance per Socket

AWS ships Graviton5 with 192 Neoverse V3 cores and 2.4x per-socket performance. M9g instances beat R8g price/performance by up to 33.6 percent amid a DRAM crunch.

Read
3 min
Words
627
Node
20nm
Operator
Amara Osei

Spec summary

  1. Graviton5 ships with 192 Neoverse V3 cores across four 48-core chiplets linked by 420 GB/sec die-to-die interconnects
  2. A Graviton5 socket delivers 2.4x the performance of Graviton4 and 25 percent more throughput than a two-chip Graviton4 NUMA node
  3. M9g instances deliver 31.9–33.6 percent better price/performance than equivalent R8g instances
  4. The chip uses a 3-nanometer process, up from 4 nanometers, and DDR5 memory at 8.8 GHz
  5. Estimated socket power is 650 watts, halving performance per watt versus Graviton4

AWS this week began shipping its Graviton5 Arm server CPU in new M9g and M9gd instances, delivering 2.4x more performance per socket than Graviton4 at roughly half the instance price of comparable X8g machines. The chip packs 192 Neoverse V3 cores — double the count of its predecessor — and lands seven months after AWS's Annapurna Labs division previewed it at re:Invent in December.

What changed from the re:Invent preview?

The block diagram AWS showed at the conference was inaccurate. It depicted a monolithic die with 96 pairs of "Poseidon" Neoverse V3 cores. The actual Graviton5 comprises four CPU chiplets, each carrying 48 V3 cores with its own memory and I/O controllers.

The design appears to build on Arm Holdings' Poseidon Compute Subsystem, cut back from 64 cores to 48 cores per block. Four die-to-die interconnects link the chiplets into a virtual processor, each running at 420 GB/sec. A full socket includes:

  • 192 Neoverse V3 cores
  • 12 DDR5 memory controllers
  • 8 PCI-Express 6.0 controllers with roughly 96 lanes, supporting CXL 3.0 memory extension

Why chiplets instead of a monolithic die?

The D2D interconnects burn significant energy, but four smaller chiplets yield far better than a monolithic design pushing against reticle limits at Taiwan Semiconductor Manufacturing Co. That saving is partially offset by the move from the 4-nanometer process used for Graviton4 to a more expensive but denser and more power-efficient 3-nanometer process.

For comparison, the Graviton4 with 96 cores carried an estimated 73 billion transistors. Graviton3 had 64 "Zeus" V1 cores; Graviton4 had 96 "Demeter" V2 cores. With Graviton4, AWS built its first two-socket NUMA machines to reach single-node performance that Graviton5 now achieves with one chip — a single Graviton5 delivers 25 percent more raw throughput than that two-chip NUMA node.

How does the cache and memory configuration stack up?

L1 and L2 cache have scaled linearly with core count, but L3 cache has grown faster. Graviton5 carries 2 MB of L2 cache per core — 384 MB total across 192 cores — plus an L3 cache of twice that size. AWS pairs the chip with the fastest DRAM available in the DDR5 form factor, running at 8.8 GHz.

The added cache, faster V3 clock speeds, DDR5 and chiplet interconnects push the socket to an estimated 650 watts. Performance per watt is therefore half that of Graviton4 — a tradeoff AWS accepts because databases and agentic AI workloads need low latency more than low heat.

What does the pricing look like?

M9g instances deliver between 31.9 percent and 33.6 percent better price/performance than the most equivalent R8g instances with the same vCPU counts. The M9gd variants with local flash deliver between 22.1 percent and 32 percent better price/performance than R8gd. Compute is cheap; memory is not:

  • M9g instances cost about half the price of X8g instances, matching their quarter-to-half memory capacity
  • Graviton5 instances carry a quarter to half the memory of their Graviton4-based predecessors
  • Savings plan pricing for M9g, M9gd and X8g has not been announced publicly
  • High-memory X8g instances remain pricey and only make sense when the memory footprint is truly needed

The memory configuration reflects the ongoing DRAM and flash crunch. For bandwidth-sensitive workloads, capacity may matter less than speed, and these instances compensate in bandwidth what they lack in capacity. Heavier-memory X9g instances are likely in the near future, but extra DRAM will carry a hefty premium at current prices.

CXL 3.0 memory extenders could mitigate some memory cost, and AWS may well have built a shared memory appliance in its Graviton5 racks to deliver that function — a more likely route than in-node PCIe 6.0 memory cards.

via The Next Platform (Source)

Filed under

  • graviton5
  • arm-neoverse-v3
  • server-cpu
  • chiplets
  • tsmc-3nm
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

249 articles

Same lot · LOT-C1C6

« Previous articleNext article »