Test report DSG-9327 · Rev F · tested October 3, 2026

AI Devices & SystemsDevice under test

Google Ships LiteRT QNN Accelerator for Qualcomm NPUs

Google's new LiteRT QNN Accelerator delegates ML inference to Qualcomm NPUs, delivering up to 100x CPU speedups and full delegation for 64 of 72 benchmark models.

Read
3 min
Words
690
Node
65nm
Operator
Elena Vasquez

Spec summary

  1. LiteRT QNN Accelerator delivers up to 100x speedup over CPU and 10x over GPU across 72 benchmark models, with 64 models delegating fully to the NPU via support for 90 LiteRT ops.
  2. On Snapdragon 8 Elite Gen 5, over 56 models run in under 5ms on the NPU versus 13 on CPU; FastVLM-0.5B achieves 0.12s time-to-first-token, 11,000+ tokens/sec prefill, and 100+ tokens/sec decode.
  3. The unified workflow removes vendor SDK handling and per-SoC targeting, supports AOT and on-device compilation, and uses Google Play AI Packs with Play for On-device AI for device-specific delivery.

Google has released the LiteRT Qualcomm AI Engine Direct (QNN) Accelerator, a new backend that routes on-device machine learning inference to the neural processing units in Qualcomm SoCs. Developed with Qualcomm, it replaces the previous TFLite QNN delegate and targets two problems at once: the fragmented NPU deployment workflow on Android and the performance ceiling of GPU-only inference.

The hardware case is straightforward. GPU compute is available on roughly 90% of Android devices, but a GPU alone bottlenecks under combined workloads — for example, a text-to-image generation model running while an ML-based segmentation processes a live camera feed. Qualcomm NPUs deliver tens of TOPS of dedicated AI compute, run in parallel with the GPU and CPU, and consume significantly less power per TOP than either. Over 80% of recent Qualcomm SoCs ship with an NPU.

Benchmark results

Google benchmarked the accelerator across 72 canonical ML models spanning vision, audio, and NLP. NPU acceleration delivers up to a 100x speedup over CPU and 10x over GPU. The accelerator supports 90 LiteRT ops, and 64 of the 72 models delegate fully to the NPU — full delegation being a critical factor for peak performance.

On the Snapdragon 8 Elite Gen 5, measured on a Xiaomi 17 Pro Max, more than 56 models complete inference in under 5ms on the NPU. Only 13 models hit that threshold on the CPU. Across 20 representative models, GPU inference reduces latency to roughly 5–70% of the CPU baseline, while the NPU brings it down to roughly 1–20%.

LLM inference: FastVLM numbers

For large language model workloads, Google benchmarked Apple's FastVLM-0.5B, a vision model designed for on-device AI, using LiteRT for both ahead-of-time (AOT) compilation and on-device NPU inference. The model uses int8 weight quantization and int16 activation quantization, which unlocks the NPU's high-speed int16 kernels. Google also added specialized NPU kernels for performance-critical transformer layers, particularly the attention mechanism.

On the Snapdragon 8 Elite Gen 5 NPU, the FastVLM integration achieves a time-to-first-token of 0.12 seconds on 1024x1024 images, over 11,000 tokens/sec for prefill, and over 100 tokens/sec for decode. Google demonstrated the setup with a live scene understanding demo that processes and describes the environment around the user.

Deployment workflow

The accelerator removes two long-standing requirements: interacting with low-level vendor-specific SDKs, and targeting individual SoC versions. LiteRT integrates with SoC compilers and runtimes behind a unified API and abstracts SoC fragmentation, so one workflow scales across supported devices. Developers can use AOT compilation or compile on the device. Pre-trained .tflite models are available from Qualcomm AI Hub.

Google recommends AOT compilation for large models, since on-device compilation increases initialization time and peak memory consumption. The process runs in a few lines of Python: developers call aot_compile against either all supported SoCs or specific targets — for example, the Snapdragon 8 Elite Gen 5 (SM8850) — then export the compiled variants into a single Google Play AI Pack. Google Play's Play for On-device AI (PODAI) mechanism delivers the correct compiled model to each user's device automatically.

On the app side, developers either copy the original .tflite file into the app's assets for on-device compilation, or integrate the AI Pack through Gradle configuration. A script fetches the QNN libraries: the NPU runtime for both workflows, plus the compiler library for on-device compilation. NPU runtime libraries attach as Gradle feature modules.

Inference itself requires a handful of Kotlin calls. Developers load a model with CompiledModel.create, pre-allocate input and output buffers, write inputs, invoke run, and read results. The runtime includes a built-in fallback mechanism: if the NPU is unavailable, LiteRT automatically uses CPU, GPU, or both as specified. AOT-compiled models support partial delegation, running unsupported subgraphs on CPU or GPU.

Google published an image segmentation sample app, an AOT compilation notebook, and documentation on the LiteRT DevSite and GitHub repository. The company credits the work to its ODML team and the Qualcomm LiteRT team, listing roughly 40 named contributors across the two groups.

via storage.googleapis.com (Original)

Filed under

  • google
  • litert
  • qualcomm
  • npu
  • on-device-ai
Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Senior reporter covering industry trends and analytics at Die Signal.

53 articles

Same lot · LOT-C1AA

« Previous article