Test report DSG-1144 · Rev C · tested October 10, 2026

Edge AI SiliconDevice under test

Arteris Proposes Direct NoC Transport For Multi-Die AI SoCs

Coherent 64-byte cache-line transactions waste bandwidth on megabyte-scale AI streams, according to Arteris, which proposes direct NoC-to-NoC transport with an invariant packet interface.

Read
4 min
Words
815
Node
65nm
Operator
Grace Kim

Spec summary

  1. Coherent die-to-die protocols transmit 64-byte cache-line transactions, each carrying addresses, transaction IDs, ordering, response tracking, and coherence metadata.
  2. Physical AI streaming transfers span hundreds of kilobytes to multiple megabytes per stream, repeating protocol overhead thousands of times per payload.
  3. Automotive AI platforms are projected to require terabytes-per-second of accelerator-to-accelerator bandwidth in future generations.
  4. Existing non-coherent approaches include AMD/Xilinx AXI Chip2Chip, Tenstorrent's Open Chiplet Atlas AXI over UCIe, and proprietary AXI-over-serial interfaces.
  5. Arteris proposes an invariant packet interface supporting multiple PHY options — UCIe, BoW, PCIe, or vendor-specific.
A Network-on-Chip (NoC) For Multi-Die Devices
Fig. AA Network-on-Chip (NoC) For Multi-Die Devices — AI-generated

Future automotive AI platforms will move terabytes of data per second between accelerator chiplets, but today's 64-byte coherent die-to-die transactions waste a growing share of that bandwidth on protocol overhead, according to Arteris.

The paper argues that coherent interconnects, designed for cache-line-sized traffic between processor cores, do not match the continuous streaming traffic generated by physical AI workloads in vehicles, robots, and industrial systems.

What problem does physical AI expose in today's die-to-die interconnects?

Physical AI workloads require simultaneous execution of multiple pipelines on a single automotive SoC. The pipelines include:

  • Camera processing
  • Radar and lidar signal processing
  • Sensor fusion
  • AI perception networks
  • Localization and mapping
  • Path planning
  • Functional safety monitoring
  • Vehicle actuator control

Rather than exchanging individual cache lines, these engines exchange continuous streams of feature maps, tensors, point clouds, video frames, and intermediate inference results. Each transfer spans hundreds of kilobytes to multiple megabytes, moving between specialized accelerators at deterministic rates.

Why are coherent protocols mismatched to AI traffic?

Coherent protocols divide communication into 64-byte transactions. Each cache line carries addresses, transaction identifiers, ordering information, response tracking, and coherence metadata.

When transferring large tensors or feature maps, this overhead repeats thousands of times. For AI workloads, payload accounts for a shrinking share of transmitted bandwidth as protocol overhead consumes a larger fraction.

Some coherent protocols have added burst optimizations, but those span only a small number of beats and cannot reduce overhead across very long streaming transfers. As accelerator bandwidth enters the terabytes-per-second range, repeated cache-line metadata becomes hard to justify.

What existing approaches fall short?

Three non-coherent die-to-die approaches ship today or have been proposed:

  • AMD/Xilinx AXI Chip2Chip, deployed in FPGAs
  • Tenstorrent's Open Chiplet Atlas AXI over UCIe (AoU)
  • Proprietary approaches that serialize AXI over high-speed serial interfaces

These extend AXI across die boundaries and work well for peripherals and control-oriented IP. They still operate at the transaction layer, however. Each NoC packet of each virtual channel must convert into AXI transactions before crossing the die boundary, then reconstruct as NoC packets on the receiving side.

This repeated conversion adds latency, buffering, and implementation complexity as systems scale to dozens of AI engines across multiple chiplets.

How does direct NoC-to-NoC transport work?

A direct NoC-to-NoC architecture lets the NoC itself span multiple dies. Native NoC packets transport directly across the die-to-die interface. The remote die becomes another portion of the same logical network.

Five technical advantages follow:

  • Streaming transfers remain packetized from producer to consumer without cache-line conversion; protocol overhead amortizes across long streams
  • Virtual channels carry multiple traffic classes — control, real-time sensor data, AI inference, bulk DMA — over the same physical link with QoS preserved
  • NoC routing, arbitration, congestion management, and QoS extend naturally across dies
  • Multiple transport protocols share one die-to-die interface: AXI, APB, AHB, or OCP
  • The die boundary becomes an additional hop, not a protocol translation point

What makes an invariant interface necessary?

Direct exposure of an implementation's internal NoC protocol creates problems. Internal packet formats evolve as NoC implementations add improved routing algorithms, new arbitration policies, additional virtual channels, enhanced QoS mechanisms, and updated packet formats.

Routing information may also depend on destination NoC topology. Without architectural stability, every NoC revision risks requiring die-to-die interface updates — workable for a single design team, impractical when teams work independently.

Arteris proposes defining an invariant architectural interface, not an exposed internal NoC. Such an interface specifies a stable packet format while allowing each chiplet's internal NoC implementation to evolve independently.

What capabilities does the invariant interface include?

The proposed interface specifies seven capabilities:

  • Support for multiple non-coherent protocols over one NoC transport layer
  • A stable packet format
  • Configurable interface widths to scale bandwidth
  • Virtual channels for multiple traffic classes
  • NoC routing semantics preserved while implementation details stay abstract
  • Isolation of future NoC enhancements from the die-to-die interface
  • PHY flexibility: UCIe, BoW, PCIe, or vendor-specific

This allows chiplets from different product generations to interoperate while each side of the boundary evolves incrementally.

Why does this matter for automotive physical AI?

Centralized vehicle compute platforms will integrate multiple AI chiplets through advanced packaging. Individual chiplets may specialize in:

  • Vision processing
  • Neural network inference
  • Radar processing
  • Lidar processing
  • CPU clusters
  • Safety and security islands
  • Memory expansion
  • Networking

The largest traffic flows often run between accelerators, not between CPUs and memory. Example pipelines include camera processing → AI perception → sensor fusion → planning → vehicle control, plus shared tensor movement between multiple neural-network accelerators.

A native packetized NoC enables higher bandwidth utilization, lower latency, and reduced buffering. The result: higher sustained throughput, lower energy per transferred bit, and greater scalability as accelerator chiplets are added — while meeting the determinism and functional safety requirements of safety-critical automotive deployments.

via hubs.ly (Original)

Filed under

  • arteris
  • network-on-chip
  • chiplets
  • die-to-die-interconnect
  • automotive-ai
Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering marketplaces and e-commerce at Die Signal.

254 articles

Same lot · LOT-C1C6

« Previous article