Test report DSG-6109 · Rev D · tested October 10, 2026

Processors & AcceleratorsDevice under test

Adreno X2: Qualcomm's Snapdragon X2 Elite GPU Doubles Compute, Falters in GPGPU

Adreno X2-90 nearly doubles Adreno X1's theoretical compute at 1.85 GHz with a 21 MB tile buffer, winning at rasterization while regressing in GPGPU workloads like FluidX3D.

Read
4 min
Words
732
Node
10nm
Operator
Grace Kim

Spec summary

  1. Adreno X2-90 runs 8 SPs in 4 slices at 1.85 GHz, up from 6 SPs at 1.5 GHz on Adreno X1, nearly doubling theoretical compute throughput.
  2. A 21 MB on-chip HPM block replaces Adreno X1's 3 MB GMEM, sized to hold a full QHD+ frame as a single tile; only 1 MB is allocatable as compute local memory.
  3. The 192-bit LPDDR5X-9523 setup delivers just over 150 GB/s, under 70% of theoretical DRAM bandwidth.
  4. In 3DMark Solar Bay, Adreno X2 beats Meteor Lake's iGPU and the Radeon 780M by more than 2x, but loses FluidX3D to the GTX 1050 3 GB.
  5. Adreno X2 adds DXR 1.1 raytracing support, with one raytracing unit per uSPTP embedding an L0 cache.

Qualcomm's Adreno X2-90, the integrated GPU in the Snapdragon X2 Elite, delivers nearly twice the theoretical compute throughput of its Adreno X1 predecessor while stepping up clock speeds from 1.5 to 1.85 GHz. Independent testing on an ASUS Zenbook A16 shows a GPU that performs strongly in rasterization and mobile-class raytracing, yet regresses against its predecessor in general-purpose compute workloads such as FluidX3D.

What changed in the architecture?

A fully enabled X2-90 part has eight Shader Processors (SPs) grouped into four slices, up from six SPs in three slices on Adreno X1. Each SP contains two Micro Shader Processor Texture Processors (uSPTPs), which carry 128 FP32 lanes arranged in what appear to be two execution unit partitions. FP16 benefits from double-rate execution. To feed the larger, faster GPU, Qualcomm scaled up cache sizes and bandwidth across the memory hierarchy.

Integer throughput improved substantially. INT32 adds appear to execute at full rate, and 32-bit integer multiplies run at half rate — comparable to Intel and AMD, and aligned with Nvidia. On Adreno X1, even basic integer operations executed at less than half rate. Like previous Adreno GPUs, Adreno X2 lacks FP64 support and hardware fused multiply-add (FMA); calling OpenCL's fma() built-in results in very poor throughput, while AMD, Intel, and Nvidia GPUs handle FMA natively without loss.

How fast is it in practice?

OpenCL testing with large dispatches shows the GPU exceeding 1 TFLOPS with FP32 adds and reaching 6 TFLOPS with multiply-adds. But it needs roughly 8x more parallelism than typical to approach those figures, and INT32 throughput barely improves over Vulkan results. Testing suggests each uSPTP partition has 6 to 8 wave slots, with 8 the more likely figure; latency more than doubles when dispatches grow from 16,384 to 18,432 workitems.

The memory hierarchy pairs 4 KB texture caches per uSPTP with 128 KB cluster caches, a 2 MB L2, an 8 MB system level cache at just over 200 ns latency, and a 192-bit LPDDR5X-9523 DRAM configuration delivering just over 150 GB/s — under 70% of theoretical. The tester's interpretation of stride-dependent latency behavior points to a virtually addressed L2 with 4 KB pages and a 16K-entry TLB placed after the L2, implying high TLB miss penalties on large-footprint accesses.

What is Adreno High Performance Memory?

The most unusual feature is a 21 MB on-chip block called Adreno High Performance Memory (HPM), seven times larger than Adreno X1's 3 MB GMEM. Qualcomm sized it to hold QHD+ frames: a 2880x1800 RGBA frame occupies 20.7 MB and should barely fit. The strategy effectively makes the whole screen one tile, eliminating tile-level rasterization inefficiencies while keeping intermediate render state on-chip.

For compute, however, only 1 MB of HPM is allocatable as OpenCL local memory — up from 384 KB on Adreno X1, but far below the usable total. Local-memory atomic throughput is also weak: 32 INT32 atomic adds per cycle across the GPU, versus 32 per WGP on AMD's RDNA 3.5.

Where does it win and lose?

Graphics results are competitive. In Cyberpunk 2077's built-in benchmark, Adreno X2 delivers twice the performance of Intel Meteor Lake's iGPU and Nvidia's GTX 1050 3 GB. In 3DMark Solar Bay it beats Meteor Lake and AMD's Radeon 780M by more than 2x and lands not far off AMD's Radeon 8060S. It also brings DXR 1.1 support to Qualcomm's GPU line.

Compute tells the opposite story. FluidX3D regresses compared to Adreno X1 even with a legacy multiply-add setting, and loses to the GTX 1050 3 GB despite having more than twice that card's compute, bandwidth, and cache capacity. In FAHBench's memory-bound nav workload, the GTX 1050 3 GB is 36% faster, versus 13.7% with the compute-bound dhfr workload. Enabling raytraced effects in Cyberpunk 2077 drops Adreno X2 well behind Meteor Lake's iGPU.

The tester's verdict: "It feels like Qualcomm took a specific set of graphics workloads and optimized for them, while AMD, Intel, and Nvidia set out to build more general purpose designs." Subjectively, the GPU is good enough not to hold back Qualcomm's laptop push, though binary translation overhead remains a larger obstacle for x86-64 gaming than the GPU itself.

via substackcdn.com (Original)

Filed under

  • qualcomm
  • adreno-x2
  • snapdragon-x2-elite
  • gpu
  • opencl
Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering marketplaces and e-commerce at Die Signal.

250 articles

Same lot · LOT-C1C6

« Previous articleNext article »