Test report DSG-1850 · Rev E · tested October 10, 2026

Processors & AcceleratorsDevice under test

Do We Still Need GPUs? Three HPC Researchers Say Maybe Not

LineShine, an all-CPU supercomputer in Shenzhen, reaches 2.74 exaflops peak theoretical and 2.2 exaflops on HPL while drawing 42.2 MW. Three leading HPC researchers argue the GPU is no longer architecturally required.

Read
4 min
Words
743
Node
10nm
Operator
Priya Raman

Spec summary

  1. LineShine hits 2.74 exaflops peak theoretical and 2.2 exaflops on HPL, scoring 52.1 gigaflops per watt while drawing 42.2 MW
  2. The LX2 CPU runs at 1.55 GHz in a 650-watt thermal envelope, likely fabricated by SMIC on a 7-nanometer process
  3. Jack Dongarra, Torsten Hoefler, and Satoshi Matsuoka co-authored the ACM paper 'Do We Still Need GPUs? Rethinking AI and Scientific Computing on Matrix-Enhanced CPUs'
  4. IBM Power10 (September 2021) and Power11 (July 2025) both ship with vector and matrix units in every core
  5. Roughly half of the Top500 machines remain CPU-only after fifteen years of GPU commercialization
Three HPC Gurus Ask: Do We Still Need GPUs?
Fig. AThree HPC Gurus Ask: Do We Still Need GPUs? — AI-generated

LineShine, the all-CPU supercomputer installed at NSC Shenzhen in China, holds the number one spot on the latest Top500 list and reaches 2.74 exaflops of peak theoretical performance using only LX2 Arm server CPUs without discrete accelerators.

A forthcoming ACM paper by Jack Dongarra (University of Tennessee / Oak Ridge National Laboratory), Torsten Hoefler (ETH Zurich / CSCS), and Satoshi Matsuoka (RIKEN / Tokyo Institute of Technology) bears the title "Do We Still Need GPUs? Rethinking AI and Scientific Computing on Matrix-Enhanced CPUs."

What does the paper actually argue?

The thesis compares two production all-CPU systems: Fujitsu's A64FX-powered Fugaku at RIKEN, in full operation since March 2021, and the recently commissioned LineShine. Both machines have hosted trillion-parameter GenAI models alongside traditional modeling and simulation.

The authors contend that GPUs became compute engines because 2010s CPUs lacked floating-point throughput, memory bandwidth, and integrated matrix engines. Those gaps, they argue, are now closing. The paper catalogues the changes:

  • Arm SVE vector extensions debuted with the A64FX in 2016
  • SVE2 vector units arrived with Armv9-A in 2019
  • Arm SME matrix units launched with Armv9-A; SME2 shipped in Armv9.4-A in 2022
  • Intel AMX matrix units reached Xeon in 2020, first deployed in Sapphire Rapids and inside the Aurora supercomputer at Argonne National Laboratory
  • AVX-512 vector units first appeared in 2016's Knights Landing Xeon Phi accelerators before migrating to Xeon proper

IBM reached production with matrix engines earlier than either competitor. Power10 shipped in September 2021 and Power11 in July 2025, both carrying vector and matrix units per core. The Telum and Telum-II mainframe chips carry their matrix engines on-die outside the cores, and the z16 and z17 mainframes add native decimal arithmetic.

Why did GPUs win originally?

In the late 2000s, GPU-accelerated nodes carried roughly a 3X price premium for 3X performance. The real win was memory bandwidth and lower energy on FP64 and FP32, not throughput per dollar. As the HPC software stack matured, the performance delta widened and price/performance improved.

The catch is the offload model. Code splits into GPU-resident parallel kernels and CPU-resident serial work, and data shuttles between devices through PCI-Express. Money flows to Nvidia as CUDA tax; memory flows to HBM. Roughly half of the Top500 machines remain CPU-only after fifteen years of GPU commercialization.

How efficient is LineShine?

The LX2 die is presumably fabricated by SMIC on a 7-nanometer process and clocks at 1.55 GHz inside a 650-watt envelope. The system draws 42.2 megawatts to deliver 2.2 exaflops on HPL, scoring 52.1 gigaflops per watt.

The paper's analysis suggests that on a TSMC 3-nanometer process, the same 2.74 exaflops might require roughly half the processors and about 25 megawatts. That would push efficiency to approximately 87 gigaflops per watt, ahead of El Capitan at Lawrence Livermore (60.9 gigaflops per watt on the June Top500) and the Nvidia Grace plus Hopper H100/H200 hybrids that lead the efficiency chart near 70 gigaflops per watt on HPL.

What changes if CPUs absorb the GPU?

The central claim is direct. "The argument is that GPUs are not fundamentally required if the CPU evolves to include the architectural features that made GPUs attractive," Dongarra, Hoefler, and Matsuoka write. "A CPU with SVE/AVX, SME/AMX, HBM, multiple precision formats, and matrix-multiply capability is no longer a conventional CPU. It is a general-purpose processor with accelerator-class numerical machinery."

IBM demonstrates the pattern today: Spyre accelerator cards deliver matrix math over PCI-Express, and IBM has signalled that a future Power12 could converge the three matrix units. AMD holds matrix capability in reserve but has not enabled it in Epyc. Nvidia, the authors imply, will eventually embed tensor cores into a future Arm server CPU and reuse the CUDA-X stack across both product lines.

The converged-workload argument runs even deeper. "Future scientific applications will not simply run simulations or train neural networks separately," the paper notes. "These workflows need both AI-style tensor throughput and traditional HPC capabilities: MPI communication, double precision, sparse solvers, adaptive algorithms, file I/O, and complex control logic. A CPU with integrated matrix acceleration may be a cleaner platform for this convergence than a system that requires constant movement between a host CPU and a discrete GPU."

The paper is available on arXiv ahead of formal ACM publication.

via linkedin.com (Original)

Filed under

  • cpus-vs-gpus
  • hpc
  • matrix-extensions
  • supercomputers
  • lineshine
Share this article:

More from Priya Raman

Priya Raman

Show full bio

Correspondent covering business strategy at Die Signal.

243 articles

Same lot · LOT-C1C6

« Previous article