Test report DSG-3840 · Rev E · tested October 10, 2026

Processors & AcceleratorsDevice under test

NVIDIA's Olympus Server Core: 10-Wide, 3.3 GHz, Near-Desktop IPC

NVIDIA's Olympus runs at 3.3 GHz yet nearly matches desktop cores in single-threaded SPEC CPU2026, with 96 KB L1D, 2 MB L2 and statically partitioned 'spatial multithreading' across 88 Vera cores.

Read
5 min
Words
1,027
Node
28nm
Operator
Amara Osei

Spec summary

  1. Olympus is a 10-wide out-of-order core running at 3.3 GHz with roughly 606 allocated register capacity
  2. SPEC CPU2026 places it slightly behind Zen 5 and ahead of Lion Cove in branch prediction accuracy
  3. Vera increases L3 capacity from 114 MB (Grace) to 164 MB and doubles per-core L2 to 2 MB
  4. A single thread achieves 88.53 GB/s from L3, implying 50-51 outstanding requests
  5. SMT causes throughput losses of 11.98% and 9.78% on some backend-bound floating point subtests
NVIDIA’s Olympus Core: Pushing Server Single Threaded Performance Boundaries
Fig. ANVIDIA’s Olympus Core: Pushing Server Single Threaded Performance Boundaries — AI-generated

NVIDIA's Olympus server core runs at just 3.3 GHz, yet a detailed microarchitectural analysis shows it nearly matches desktop-class single-threaded performance from Intel and AMD while dominating Arm's Neoverse server competition.

Olympus is a 10-wide out-of-order core in NVIDIA's Vera CPU, shipping with 88 cores. It pursues performance per clock rather than clock speed, with very large out-of-order structures and a simultaneous multithreading (SMT) implementation NVIDIA calls "spatial multithreading."

Server chips traditionally trailed client parts in single-threaded performance because higher core counts leave less power per core, and complex interconnects add latency. AMD narrowed that gap over recent generations by pushing frequencies. Olympus attacks it from the opposite direction: maximum work per cycle at modest clocks.

What does the core look like?

The general layout resembles Arm's Cortex X925. Both use a semi-distributed scheduler with a similar execution unit layout, and both target high performance without AMD- or Intel-level clock speeds. Olympus goes further with larger out-of-order structures than X925, plus a non-scheduling queue in front of the floating point schedulers.

Key frontend specifications:

  • 64 KB, 4-way set associative instruction cache delivering 128 bytes per cycle into a 48-instruction decode queue
  • 64-entry fully associative instruction TLB
  • 32B/cycle instruction fetch from L2 for large footprints, higher single-thread L2 code throughput than Zen 5 or Cortex X925
  • 16K-entry BTB with 4-cycle latency for footprints beyond 48 KB

The backend has massive register files and queues throughout. NOP testing shows the core can keep more than 1000 instructions in flight, though a practical cap of roughly 606 allocated architectural registers limits the real figure. Eight integer ALUs sit in pairs, each fed by a scheduling queue of approximately 25 entries. A six-pipe FPU uses a non-scheduling queue feeding smaller scheduling queues, where Arm uses three very large ones.

Memory access runs down four pipelines — all four handle loads, two handle stores. The L1 data cache is 96 KB, 6-way set associative, with a nominal 4-cycle latency; simple tests measure 2-3 cycles, possibly due to value prediction.

How does branch prediction compare?

In SPEC CPU2026, Olympus lands slightly behind AMD's Zen 5 and slightly ahead of Intel's Lion Cove in prediction accuracy, with individual wins and losses across workloads. Its direction predictor tracks roughly 5-6K global history patterns with a couple of branches in play, versus 16-24K for X925 — or 48K versus over 64K patterns with 512 branches.

Like Zen 5, Olympus sustains two taken branches per cycle. Unlike Zen 5 and X925, that capability ties to instruction footprint rather than branch count: footprints beyond 48 KB break it.

NVIDIA and AMD both use ahead indexing, where one lookup predicts the next two branches. NVIDIA's Vera whitepaper claims: "Olympus implements advanced predictors with robust ahead pipelining mechanisms delivering up to 2.3x higher predictions per cycle vs the competition." AMD's Zen 5 Optimization Guide states: "Predicting with BTB pairs allows two fetches to be predicted in one prediction cycle." Validating the 2.3x claim directly is difficult — neither architecture exposes hardware counters for prediction counts — and Zen 5's higher clocks narrow the effective gap in branch completion rate.

How does the cache hierarchy stack up?

Vera's caching strategy improves substantially on NVIDIA's prior Grace CPU:

  • 96 KB L1D, versus the 32-64 KB typical of AMD, Arm and Intel server cores
  • 2 MB 8-way L2 at 10-cycle load-to-use latency, doubling Grace's 1 MB per-core L2
  • L3 grown from 114 MB to 164 MB, though latency exceeds 120 cycles

Aggressive prefetching offsets the L3 latency. A single-thread linear read achieves 88.53 GB/s from L3; applying Little's Law yields 50-51 outstanding L3 requests, versus 35-36 on Grace/GH200. Lemire's memory level parallelism test shows 18-19 demand L1D misses in flight.

The 112-entry fully associative data TLB is backed by a unified 3K-entry L2 TLB, against Zen 5's 96-entry DTLB and 4K-entry L2 DTLB. One caveat: the test system's kernel used 64 KB pages, while all comparison results used 4 KB pages, complicating direct SPEC comparisons.

What is 'spatial multithreading'?

NVIDIA statically partitions nearly every core resource — including the L1D, L2 caches, schedulers and even execution throughput — so each thread gets half when the sibling runs, even if that sibling makes no memory accesses. AMD's Zen 5 flexibly shares most structures with watermarking instead. The result behaves like splitting Olympus into two 5-wide cores, which simplifies fairness guarantees and design.

SMT gains on SPEC CPU2026 integer workloads match Zen 5's, with low-IPC, frontend-bound tests like 723.llvm seeing massive uplift. But backend-bound floating point tests expose weaknesses: throughput drops of 11.98% and 9.78% in specific subtests, where Zen 4 loses only 1-3%. In 709.cactus, L2 data MPKI rises from 3.58 to 6.52 with two threads, and partitioned backend resources struggle against high L3 latency.

Where does it land on SPEC?

SPEC CPU2026 places Olympus not far off the Intel Core Ultra 9 285K and Ryzen 7 9800X3D in single-threaded terms, and well ahead of Neoverse V2 and N2. IPC often exceeds Zen 5's, even where Zen 5 delivers higher absolute performance. The large loss in 706.stockfish_r stems from instruction set differences — Olympus executes roughly twice as many instructions as Zen 5 there.

SPEC CPU2017 data shows Olympus beating several server cores including a Zen 5 server implementation, and drawing even with AMD's top-end Zen 4 desktop chip.

Vera reverses the Arm server dynamic: NVIDIA now holds fewer cores with stronger per-core throughput, while AMD ships more cores with less single-threaded performance. The page-size discrepancy leaves some wiggle room in the SPEC results, but likely narrows rather than reverses the rankings. Against Neoverse cores — including Grace, which cannot run a second thread per core at all — Olympus widens the per-core gap in nearly every case.

via substackcdn.com (Original)

Filed under

  • nvidia
  • olympus
  • vera
  • microarchitecture
  • server-cpu
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

249 articles

Same lot · LOT-C1C6

« Previous articleNext article »