Test report DSG-1419 · Rev B · tested October 10, 2026

Memory & StorageDevice under test

Hot Chips 2026: High Bandwidth Flash targets AI memory crunch

At Hot Chips 2026, Anurag Agarwal and Radhakrishna Giduthuri outlined how High Bandwidth Flash could ease HBM capacity limits for ML workloads — if software adapts.

Read
3 min
Words
585
Node
20nm
Operator
Amara Osei

Spec summary

  1. HBF cubes place SSD-type flash on the processor package in an HBM-like form factor to deliver higher capacity than HBM.
  2. No HBF products exist yet; the Hot Chips 2026 talk relies on simulations and projections.
  3. HBF requires large, aligned block accesses — a single-byte write can force a full 64 KB read-modify-write cycle.
  4. vLLM use cases include storing MoE experts and KV cache in HBF with DMA transfers into HBM.
  5. HBF wins on cost per capacity but loses to HBM on cost per bandwidth, favoring workloads below bandwidth limits.
Hot Chips 2026: Applying High Bandwidth Flash (HBF)
Fig. AHot Chips 2026: Applying High Bandwidth Flash (HBF) — AI-generated

No HBF product exists today, yet a Hot Chips 2026 tutorials-day talk from Anurag Agarwal and Radhakrishna Giduthuri laid out how High Bandwidth Flash could reshape memory hierarchies for machine learning. The proposal: package SSD-grade flash in HBM-style cubes that sit on the same substrate as the compute chip, trading bandwidth for far higher capacity.

HBF uses the same NAND flash technology found in SSDs, but it is not another Optane-style memory pool. Despite the HBM-like form factor, it behaves like an SSD integrated onto the processor.

What does HBF actually require from software?

Access works through DMA, not ordinary memory operations. Key constraints:

  • All accesses must occur in large, aligned chunks, as with a mass storage device rather than system memory.
  • Host software takes on SSD-controller duties such as write leveling and data retention management.
  • Modifying a single byte can mean reading a 64 KB block into DRAM, changing it, and writing the entire block back to flash.
  • Existing DRAM-oriented frameworks need massive changes to exploit HBF — and switching frameworks means redoing that work.

The analogy, as the source analysis puts it, is a low-level disk API: "like using FILE_FLAG_NO_BUFFERING in Windows or O_DIRECT in Linux." HBF is not plug-and-play.

How could vLLM use it?

Giduthuri used vLLM as a case study. vLLM already explores ways to cut VRAM usage, such as holding model weights in pinned CPU memory — an approach that fails on HBF because HBF lacks fine-grained random access. Other paths look more viable:

  • MoE experts in HBF. Software DMA-moves only active experts into HBM as needed.
  • KV cache in HBF. This works best with sparse attention that reads only a subset of tokens per step, letting most of the cache sit "cold" in flash. One caveat: top-k reads are scattered while HBF prefers sequential access; DMA-ing top-k rows into DRAM on demand may resolve this.

Can HBF cut cross-GPU traffic?

Large models are typically sharded across GPUs, and cross-device scatter/gather operations often become a bigger performance barrier than compute throughput or memory bandwidth. HBF capacity lets designers replicate more model weights across GPUs instead. DMA off flash is not cheap, but it is cheaper than going off-device.

When does the cost math favor HBF?

Agarwal framed the economics plainly. HBF works when a workload does not hit its bandwidth limits — smaller models, smaller batch sizes. Once a workload becomes bandwidth-bound, the equation flips: both cost per capacity and cost per bandwidth factor into final cost, and HBF wins on capacity while losing to HBM on bandwidth. Using HBM to cache hot experts could help, but the caching must work out well, or HBF bandwidth will damage the cost-per-token result.

Is HBF easier than just streaming from an SSD?

The source analysis raises a blunt counterpoint: the engineering effort to leverage HBF looks close to what it would take to simply reduce DRAM usage by streaming model weights off a conventional SSD — and the SSD route appears easier. Kernel buffering abstracts away block alignment, allows arbitrary seeks and byte-level reads/writes, and naturally caches against flash inefficiencies. Existing SSD weight-streaming projects might transfer to HBF, or the software challenges might block adoption entirely.

The verdict waits on shipping silicon. HBF may relieve the DRAM capacity problem to some extent, but the software burden is substantial, and whether that trade pays off remains unproven until products reach the market.

via github.com (Original)

Filed under

  • hbf
  • hot-chips-2026
  • ai-memory
  • hbm
  • vllm
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

254 articles

Same lot · LOT-C1C6

« Previous article