Test report DSG-6688 · Rev E · tested September 29, 2026

Processors & AcceleratorsDevice under test

Arm's C2-Ultra, G2-Ultra NX and Neoverse CSS N4: The Numbers Behind the Claims

Arm's C2-Ultra gains 7% peak IPC over C1-Ultra; G2-Ultra NX adds matrix accelerators and cuts RT DRAM traffic 13%; CSS N4 spans 8–128 cores at up to 3.8 GHz.

Read
4 min
Words
865
Node
20nm
Operator
Grace Kim

Spec summary

  1. C2-Ultra's maximum IPC increase over C1-Ultra is 7%; the average uplift drops to 3.2% once an 8.5% clock advantage is factored out, and still includes an extra 1 MB of L2.
  2. G2-Ultra NX's optional matrix accelerator delivers 1,024 INT8 MACs per clock but lacks FP8 and BF16 support; it adds about 21% to shader core area and requires six NX cores for Ultra branding.
  3. Neoverse CSS N4 supports 8–128 cores per die at up to 3.8 GHz, up to 256 MB shared cache, DDR5/LPDDR6, and 128 lanes of PCIe Gen 6/7 with CXL 4.0.
Arm’s C2-Ultra, G2-Ultra NX, and CSS N4 IP
Fig. AArm’s C2-Ultra, G2-Ultra NX, and CSS N4 IP — AI-generated

Arm's latest IP announcements span mobile and datacenter: the C2-Ultra CPU core, the G2-Ultra NX GPU, and the Neoverse CSS N4 platform. The G2-Ultra NX already ships in Xiaomi's XRING O3, launched August 24. The disclosures, however, mix genuine microarchitectural changes with performance figures that fold in clock speed, cache size, and memory bandwidth gains.

C2-Ultra: Iterative Changes, Complicated Claims

Compared with the two-generation-old Cortex X925, C2-Ultra retains the same overall layout: 10-wide decode, 8 simple ALUs, 6 FP lanes, 3 branch ports, and a 4-load/2-store configuration. The gains come from targeted changes to the branch predictor and out-of-order buffers. Arm stayed vague on specifics — the company did not disclose whether it modified BTB sizes, return stacks, or the prediction algorithm. It also cites a larger execution window and improved speculation versus C1-Ultra, without explicit details. The net effect: C2-Ultra spends less time waiting for data.

Arm claims a 15% peak performance uplift over C1-Ultra and a 12% average gain in traditional benchmarks. The endnotes carry critical context. First, the numbers come from FPGA simulations. Second, the C2-Ultra platform ran at an 8.5% higher clock and with a 3 MB L2 cache — a configuration C1-Ultra does support and already ships with in Xiaomi's XRING O3 and Samsung's Exynos 2600. Arm states the memory subsystem delivers nearly twice the bandwidth in CPU benchmarks but does not factor into the claimed uplift.

After factoring out the clock increase, the average uplift drops to 3.2%, and that figure still includes the extra 1 MB of L2, which benefits some workloads — Geekbench 6, for instance, gains from a larger L2 that catches most L1 misses. Following a clarification Arm issued on September 8, the maximum IPC increase over C1-Ultra, using the endnote parameters, is 7%. On the tested workloads, C2-Ultra appears to improve performance per clock only modestly.

Arm also claims 38% lower power versus C1-Ultra, but that number includes node and implementation improvements, leaving the microarchitecture's contribution unclear. The company additionally announced C2-Nano and C2-Pro; both reuse the C1-Nano and C1-Pro microarchitectures.

G2-Ultra NX: Registers, Ray Tracing, Matrix Units

Arm calls Mali G2-Ultra NX "the largest re-architecting of the GPU IP in 7 generations." Raw throughput is unchanged: each shader core still has 128 FMA units (256 FP32 FLOPs per clock, 512 FP16), and the maximum of 24 shader cores per GPU carries over.

The rework targets registers and memory traffic. A warp can now access 128 registers, up from 64, allocation granularity improves to 16 registers, and the register file grows 25%. A more compact triangle structure in the RT units eliminates redundant data and cuts DRAM traffic by 13%. Arm also brings Opacity Micromaps, a desktop GPU feature, to mobile.

The other desktop-class addition is a matrix accelerator inside the shader core, capable of 1,024 INT8 or 512 INT16 MACs per clock, and able to run at up to twice the Execution Engines' clock. Notably, it supports neither FP8 nor BF16, and BF16 is absent from the standard ALUs as well. Arm does not require every shader core to carry an accelerator; the Ultra branding requires a minimum of six "NX cores." Xiaomi's XRING O3 implements the unit on half its shader cores. A shader core with the matrix unit measures about 1.88 mm² versus 1.55 mm² without it — roughly a 21% area penalty.

The matrix units enable Arm's Neural Super Sampling (NSS) upscaling, and the G2 generation adds frame rate upscaling. Arm claims up to 14% improvement in games and 24% in ray tracing benchmarks — but the G2-Ultra NX runs about 11% higher clocks than G2-Ultra in those comparisons, suggesting limited gains in non-RT titles.

Neoverse CSS N4

On the server side, Neoverse CSS N4 will become available to partners. Arm's presentation contained a single slide; further details emerged only in response to questions:

  • CPU: 8–128 Neoverse N4 cores per die, the widest core-count range in a Neoverse CSS, at up to 3.8 GHz
  • Multi-chiplet and multi-socket scaling beyond a single die
  • 64 KB instruction and 64 KB data L1 per core; up to 2 MB private L2 per core; up to 256 MB shared system-level cache per die
  • DDR5 or LPDDR6 memory support
  • Up to 128 lanes of PCIe Gen 6/7 and CXL 4.0
  • Arm chip-to-chip interconnect with UCIe or partner-specific PHY support

Open questions remain, including which core N4 uses, single-die memory bus width, and single-die PCIe lane counts.

Thin Disclosure

The technical depth of this announcement falls short of Arm's prior briefings. C2-Ultra's headline numbers combine core changes with higher clocks, more L2, and more bandwidth; power figures blend process gains with architecture. G2-Ultra NX fares better on register allocation and RT details, but the "seven generations" framing set expectations the disclosure did not meet. Arm answered questions when asked — the problem is that those questions were necessary at all.

via substackcdn.com (Original)

Filed under

  • arm
  • c2-ultra
  • g2-ultra-nx
  • neoverse
  • css-n4
Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering marketplaces and e-commerce at Die Signal.

46 articles

Same lot · LOT-C1BA

Next article »