Test report DSG-4937 · Rev D · tested September 30, 2026

Edge AI SiliconDevice under test

Qualcomm Teases Hexagon NPU Upgrade for On-Device 30B-Parameter AI Agents

Qualcomm's next Hexagon NPU adds an Element Accelerator, larger shared memory, and MoE support to run 30B-parameter AI agents fully on-device.

Read
3 min
Words
620
Node
65nm
Operator
Marcus Bennett

Spec summary

  1. The upcoming Hexagon NPU supports Mixture-of-Experts models with up to 30 billion parameters running entirely on-device.
  2. A new Element Accelerator sits alongside the tensor, vector, and scalar units and speeds up transformer inference without added power cost.
  3. Qualcomm claims 50% higher pre-fill performance for INT4-precision models, alongside faster decoding throughput and speculative decoding.
Qualcomm's next Snapdragon flagship chip wants your AI agents to stay on-device - Android Authority
Fig. AQualcomm's next Snapdragon flagship chip wants your AI agents to stay on-device - Android Authority — AI-generated

Qualcomm has disclosed the first details of the next-generation Hexagon NPU that will ship in its upcoming Snapdragon Elite chips, and the focus is unambiguous: on-device agentic AI. The company says the new NPU can run Mixture-of-Experts (MoE) models with up to 30 billion parameters locally, alongside two architectural changes aimed at making multi-step AI agents faster and more power-efficient.

Two headline changes

Qualcomm highlights two key elements of the new Hexagon design. The first is a new component called the "Element Accelerator," which now sits alongside the NPU's scalar, vector, and tensor units. According to Qualcomm, it is designed to speed up transformer inference, "helping agents respond faster, reason more efficiently, and deliver richer experiences" — without adding power cost. The company did not disclose the accelerator's exact bandwidth.

The second change is "much larger shared memory." The four accelerator units — tensor, vector, scalar, and the new Element — now share a bigger memory pool, which Qualcomm says reduces dependence on the phone's main RAM module when exchanging data from the KV-cache.

In practical terms, the NPU can hold more context in fast local memory. That headroom matters most for agentic workloads, where an AI assistant must track state across multiple reasoning steps rather than answer a single query and forget everything.

MoE support brings 30B-parameter models on-device

Qualcomm also confirmed the new Hexagon NPU will support new AI architecture built specifically for Mixture-of-Experts models. MoE systems route each query through specialized "expert" sub-models based on the input, activating only a fraction of a model's total parameters at any given time. This routing cuts compute and memory bandwidth demands dramatically, and it is the mechanism that lets Qualcomm claim local execution for models with up to 30 billion parameters on a phone SoC.

That figure deserves scrutiny. A 30B-parameter model running locally does not mean all 30 billion parameters are active for every token; MoE's advantage is precisely that it activates a small subset. Still, the claim points to on-device capabilities that were server-side territory until recently.

Performance claims: read the fine print

Qualcomm quantifies the gains as "50% higher pre-fill performance, faster decoding throughput, enhanced speculative decoding, and higher overall tokens per second."

One caveat applies. The 50% figure covers pre-fill performance specifically, and the claims apply to models running at INT4 precision. It is not a blanket 50% increase in AI performance across the board. Taken together, Qualcomm says the enhancements help models deliver faster responses and complete multi-step tasks more quickly.

Why agentic workflows are the target

The strategic through-line is local processing. Qualcomm frames the combination of the Element Accelerator, larger shared memory, and MoE support as making agentic workflows — not just conventional one-shot AI queries — "more efficient and responsive." An agent that plans, calls tools, and iterates over several steps needs sustained throughput, large context retention, and low memory latency. Each of the three disclosed changes maps directly onto one of those requirements.

The disclosure follows earlier teasers covering CPU and GPU upgrades for the upcoming Snapdragon Elite chips, making this the third major block of the platform Qualcomm has detailed so far.

What remains undisclosed

Several numbers are still missing. Qualcomm has not published exact shared-memory capacity, accelerator bandwidth figures, or comparative benchmarks against the current Hexagon NPU in the Snapdragon 8 Elite Gen 5. The company also has not named the chips outright, though the upcoming silicon is expected to appear in the Snapdragon 8 Elite Gen 6 family based on prior leaks.

Expect fuller specifications, benchmark data, and device-partner announcements when Qualcomm formally launches the platform.

via androidauthority.com (Original)

Filed under

  • qualcomm
  • hexagon-npu
  • snapdragon
  • on-device-ai
  • moe
Share this article:

More from Marcus Bennett

Marcus Bennett

Show full bio

News editor covering marketplaces and e-commerce at Die Signal.

58 articles

Same lot · LOT-C1D0

« Previous articleNext article »