Test report DSG-2514 · Rev E · tested October 10, 2026
Processors & AcceleratorsDevice under test
Vera silicon outperforms NVIDIA's own whitepaper arguments
Phoronix measured Vera at 10% above an EPYC 9575F and 1.55x a Xeon 6980P in geomean, yet NVIDIA's 45-page whitepaper on the 88-core Arm CPU reaches well past the data. A 3× memory claim collapses to 1.9× under independent Turin testing.
- Read
- 5 min
- Words
- 910
- Node
- 10nm
- Operator
- Elena Vasquez
Spec summary
- Phoronix in May measured Vera at 10% higher geomean than a 5 GHz EPYC 9575F, 1.55x a Xeon 6980P, and 1.63x NVIDIA's own Grace
- Vera packs an 88-core monolithic Olympus die, 1.2 TB/s of LPDDR5X memory bandwidth, and a 164 MB shared last-level cache
- Estimated SPECrate 2026 Integer Base totals reach 925 for two Vera sockets versus 898 for two EPYC 9755 sockets, a 3.0% system-throughput advantage
- Independent testing measured ~570 GB/s from an EPYC 9755 on DDR5-6400, cutting NVIDIA's 3× memory-bandwidth claim to roughly 1.9×
- AMD launched 6th Gen EPYC two days after the whitepaper's July 21 publication, with the 96-core 9686F hitting 1,638 GB/s per socket on MRDIMM-12800

Vera's 88-core monolithic die delivered a 10% geomean lead over a 5 GHz EPYC 9575F and 1.55x a Xeon 6980P in independent Phoronix testing published in May, yet the 45-page whitepaper NVIDIA released to explain the new Arm server CPU reaches well beyond what the benchmark numbers support.
The chip itself pairs eight LPDDR5X interfaces at 1.2 TB/s with a 164 MB shared last-level cache and a new Olympus core: a 10-wide Arm v9.2 design with value prediction, a graph-aware prefetcher, 2 MB of private L2 per core, a 96 KB L1 data cache, and a 3.4 TB/s coherency fabric.
Many of those features have parallels on shipping parts. AMD shipped limited floating-point value prediction in Family 17h, Intel has run a Data-Dependent Prefetcher since at least 2022, and perceptron branch predictors trace back to AMD's 2012 Piledriver microarchitecture.
Is Spatial Multithreading really different from SMT?
Figure 5 in the whitepaper contrasts "Traditional SMT (x86)" with NVIDIA's Spatial Multithreading, depicting the x86 pipeline alternating between two threads while Vera statically partitions its two hardware threads. The caption claims Spatial Multithreading avoids "opportunistic time-sharing."
The diagram omits per-cycle thread selection in fetch and decode, and the thread-agnostic behavior of execute and memory stages that designers have used since Intel's Pentium 4 HT shipped in 2002. Both designs can hide single-thread stalls; static partitioning wastes whatever the second thread cannot consume.
One undisclosed cost: an Olympus core takes 10,000 cycles to return to single-thread mode after a sibling thread exits, per a Linux kernel patch posted to lore.kernel.org. Software writers must weight that penalty before launching a second thread on a core.
Does the 32-NUMA-node figure represent a real comparison?
NVIDIA argues that a large two-socket x86 server can expose "as many as 32 NUMA domains," against Vera's one per socket. The 32-figure is real on many-chiplet EPYC, but it sits at the end of a configurable spectrum.
AMD's EPYC 9005 tuning guide lists NPS4, NPS2, NPS1, and even NPS0 modes, plus an opt-in "LLC as NUMA" setting. The whitepaper presents an extreme configuration as though it were the typical user experience.
What changed when AMD shipped 6th Gen EPYC?
NVIDIA published its blog and whitepaper on July 21. Two days later AMD launched 6th Gen EPYC. The 96-core EPYC 9686F with 16 memory channels supports DDR5-8000 or MRDIMM-12800, reaching 1,024 to 1,638 GB/s per socket. At the top MRDIMM speed, the new EPYC platform exceeds Vera's theoretical 1.2 TB/s on total bandwidth.
That timing undercuts the "3× more memory bandwidth than the latest x86 CPU" claim the whitepaper leans on.
How big is the memory advantage really?
The whitepaper showed Vera reaching roughly 1.1 TB/s in a loaded-latency plot against a dual-socket EPYC 9755 topping out near 400 GB/s, with 12.7 GB/s per core against 3.1 GB/s per core. Those Turin numbers do not match independent test runs.
Independent testing measured about 570 GB/s on a 12-channel DDR5-6400 EPYC 9755, roughly 93% of AMD's 614 GB/s theoretical ceiling. Recalculated against Vera's 1.1 TB/s:
- Total bandwidth: ~1.9× (not the 2.7× the paper shows)
- Per-core bandwidth: ~2.8× (not 4.1×)
- Versus the 9575F AI-head-node SKU: ~1.4×, or about a 40% lead
Vera still wins because its eight LPDDR5X channels deliver roughly twice the peak bandwidth of one Turin socket, not because chiplets lose bandwidth.
Why does the "agentic" framing matter?
NVIDIA labels four SPEC CPU 2026 integer workloads (CPython, GCC, LLVM, Cppcheck) as "agentic benchmarks." SPEC itself describes them as a Python interpreter, two compilers, and a static analyzer: legitimate CPU programs, not end-to-end agent runs.
The whitepaper does correctly tag the figures as estimates. Figure 15 reports a 1.7x-1.8x per-core advantage on the selected components. The full estimated SPECrate 2026 Integer Base totals reach 925 for two Vera sockets against 898 for two EPYC 9755 sockets, a 3.0% system-throughput advantage. The 176 vs. 256 physical-core count drives most of the per-core delta.
Can a pictogram benchmark a CPU?
Figure 24, "Vera drives 1.8x for RL training," shows a row of completed-task squares with no model, environment, allocation, framework, batch size, power data, repetition count, or error bars.
The earlier PageRank chart claims a 2.6x advantage and near-linear Vera scaling, but GAP Benchmark Suite variables are omitted and the scaling plot stops at 32 of 88 cores. Without the inputs the result is not reproducible.
The contrast with x86 is set up rather than demonstrated. NVIDIA's quoted pitch reads: "By reducing resource interference between threads, Spatial Multithreading improves determinism, isolation, and quality of service compared to traditional SMT approaches. The result is a CPU architecture that can run large numbers of concurrent agent tasks while maintaining more consistent latency and throughput." Quality of service matters for the target market, but the sentence never claims a throughput win.
Independent reviewers with production hardware, unrestricted power and frequency readouts, and on/off SMT tests will resolve what the Olympus core actually does. The 45-page document argues a case the silicon does not need.
via amd.com (Original)
More from Elena Vasquez
Show full bio
Senior reporter covering industry trends and analytics at Die Signal.
251 articles
Same lot · LOT-C1C6
- DSG-468765nmAMD's Instinct MI455X Pairs 432 GB HBM4 with 72-GPU Helios Rack
- DSG-132728nmAMD Outlines Venice, MI455X, and Helios Roadmaps at Advancing AI 2026
- DSG-668820nmArm's C2-Ultra, G2-Ultra NX and Neoverse CSS N4: The Numbers Behind the Claims
- DSG-424228nmIntel's Diamond Rapids jumps to 256 cores, takes server fight to AMD Venice
- DSG-124314nmAMD EPYC 9006: Mapping Venice Silicon to Agentic AI Workloads