Test report DSG-9327 · Rev F · tested October 3, 2026
AI Devices & SystemsDevice under test
Google Ships LiteRT QNN Accelerator for Qualcomm NPUs
Google's new LiteRT QNN Accelerator delegates ML inference to Qualcomm NPUs, delivering up to 100x CPU speedups and full delegation for 64 of 72 benchmark models.
- Read
- 3 min
- Words
- 690
- Node
- 65nm
- Operator
- Elena Vasquez
Spec summary
- LiteRT QNN Accelerator delivers up to 100x speedup over CPU and 10x over GPU across 72 benchmark models, with 64 models delegating fully to the NPU via support for 90 LiteRT ops.
- On Snapdragon 8 Elite Gen 5, over 56 models run in under 5ms on the NPU versus 13 on CPU; FastVLM-0.5B achieves 0.12s time-to-first-token, 11,000+ tokens/sec prefill, and 100+ tokens/sec decode.
- The unified workflow removes vendor SDK handling and per-SoC targeting, supports AOT and on-device compilation, and uses Google Play AI Packs with Play for On-device AI for device-specific delivery.
Google has released the LiteRT Qualcomm AI Engine Direct (QNN) Accelerator, a new backend that routes on-device machine learning inference to the neural processing units in Qualcomm SoCs. Developed with Qualcomm, it replaces the previous TFLite QNN delegate and targets two problems at once: the fragmented NPU deployment workflow on Android and the performance ceiling of GPU-only inference.
The hardware case is straightforward. GPU compute is available on roughly 90% of Android devices, but a GPU alone bottlenecks under combined workloads — for example, a text-to-image generation model running while an ML-based segmentation processes a live camera feed. Qualcomm NPUs deliver tens of TOPS of dedicated AI compute, run in parallel with the GPU and CPU, and consume significantly less power per TOP than either. Over 80% of recent Qualcomm SoCs ship with an NPU.
Benchmark results
Google benchmarked the accelerator across 72 canonical ML models spanning vision, audio, and NLP. NPU acceleration delivers up to a 100x speedup over CPU and 10x over GPU. The accelerator supports 90 LiteRT ops, and 64 of the 72 models delegate fully to the NPU — full delegation being a critical factor for peak performance.
On the Snapdragon 8 Elite Gen 5, measured on a Xiaomi 17 Pro Max, more than 56 models complete inference in under 5ms on the NPU. Only 13 models hit that threshold on the CPU. Across 20 representative models, GPU inference reduces latency to roughly 5–70% of the CPU baseline, while the NPU brings it down to roughly 1–20%.
LLM inference: FastVLM numbers
For large language model workloads, Google benchmarked Apple's FastVLM-0.5B, a vision model designed for on-device AI, using LiteRT for both ahead-of-time (AOT) compilation and on-device NPU inference. The model uses int8 weight quantization and int16 activation quantization, which unlocks the NPU's high-speed int16 kernels. Google also added specialized NPU kernels for performance-critical transformer layers, particularly the attention mechanism.
On the Snapdragon 8 Elite Gen 5 NPU, the FastVLM integration achieves a time-to-first-token of 0.12 seconds on 1024x1024 images, over 11,000 tokens/sec for prefill, and over 100 tokens/sec for decode. Google demonstrated the setup with a live scene understanding demo that processes and describes the environment around the user.
Deployment workflow
The accelerator removes two long-standing requirements: interacting with low-level vendor-specific SDKs, and targeting individual SoC versions. LiteRT integrates with SoC compilers and runtimes behind a unified API and abstracts SoC fragmentation, so one workflow scales across supported devices. Developers can use AOT compilation or compile on the device. Pre-trained .tflite models are available from Qualcomm AI Hub.
Google recommends AOT compilation for large models, since on-device compilation increases initialization time and peak memory consumption. The process runs in a few lines of Python: developers call aot_compile against either all supported SoCs or specific targets — for example, the Snapdragon 8 Elite Gen 5 (SM8850) — then export the compiled variants into a single Google Play AI Pack. Google Play's Play for On-device AI (PODAI) mechanism delivers the correct compiled model to each user's device automatically.
On the app side, developers either copy the original .tflite file into the app's assets for on-device compilation, or integrate the AI Pack through Gradle configuration. A script fetches the QNN libraries: the NPU runtime for both workflows, plus the compiler library for on-device compilation. NPU runtime libraries attach as Gradle feature modules.
Inference itself requires a handful of Kotlin calls. Developers load a model with CompiledModel.create, pre-allocate input and output buffers, write inputs, invoke run, and read results. The runtime includes a built-in fallback mechanism: if the NPU is unavailable, LiteRT automatically uses CPU, GPU, or both as specified. AOT-compiled models support partial delegation, running unsupported subgraphs on CPU or GPU.
Google published an image segmentation sample app, an AOT compilation notebook, and documentation on the LiteRT DevSite and GitHub repository. The company credits the work to its ODML team and the Qualcomm LiteRT team, listing roughly 40 named contributors across the two groups.
via storage.googleapis.com (Original)
More from Elena Vasquez
Show full bio
Senior reporter covering industry trends and analytics at Die Signal.
53 articles
Same lot · LOT-C1AA
- DSG-637710nmASUS Ascent QN10: Snapdragon X2 Elite Mini PC With 80 TOPS NPU
- DSG-712865nmQualcomm Redesigns Hexagon NPU for On-Device Agentic AI
- DSG-493765nmQualcomm Teases Hexagon NPU Upgrade for On-Device 30B-Parameter AI Agents
- DSG-586145nmTI MSPM0G5187 MCUs Pair TinyEngine NPU with Cortex-M0+ for Edge AI
- DSG-28953nmNVIDIA Unveils RTX Spark AI Superchip