Test report DSG-8310 · Rev A · tested October 10, 2026

Processors & AcceleratorsDevice under test

AMD Buys Taalas to Add Model-Specific Silicon to Instinct Roadmap

AMD acquired AI inference startup Taalas this week for an undisclosed sum, the second GPU vendor in months to fold a model-specific decode accelerator into its stack after Nvidia's $20B Groq deal.

Read
3 min
Words
558
Node
28nm
Operator
Amara Osei

Spec summary

  1. AMD acquired AI inference startup Taalas this week for an undisclosed sum
  2. Nvidia paid $20 billion for Groq's engineering team and technology in December
  3. Cerebras carries a $50.9 billion market capitalization on under $200 million in quarterly revenue
  4. Taalas HC1 chip supports models up to 8 billion parameters; HC2 next generation targets 20 billion
  5. Customizing a Taalas HC chiplet costs about 1/100th the price of training a new generative AI model
With Taalas, AMD Can Bake AI Inference Directly Into Its Chippery
Fig. AWith Taalas, AMD Can Bake AI Inference Directly Into Its Chippery — AI-generated

AMD confirmed this week the acquisition of AI inference chip startup Taalas for an undisclosed sum, the second GPU vendor in four months to fold a model-specific accelerator startup into its stack.

In December, Nvidia paid $20 billion for Groq's engineering team and technology. AMD's deal buys hardcoded-model decode silicon it can pair with its Instinct GPUs.

Why split inference across two silicon types?

Jensen Huang, Nvidia's chief executive officer and co-founder, used his GTC 2026 keynote in March to argue that agentic AI workloads need latency profiles GPUs cannot deliver alone.

Huang's charts showed that an eight-GPU node built from Grace CG100 CPUs and Hopper H100 GPUs handled about 100 tokens per second per user before throughput collapsed. The Grace-Blackwell NVL72, with 18 Grace CPUs and 36 Blackwell B300 GPUs, extended interactivity to 200 TPS per user. The upcoming Vera-Rubin NVL72, with Vera CV100 CPUs and Rubin R200 GPUs, targets 400 TPS per user at roughly 10X the tokens-per-second-per-megawatt efficiency of Grace-Hopper.

At the 400 TPS Premium interactivity level, Grace-Blackwell's TPS-per-megawatt drops close to zero. Pairing it with Groq-class SRAM-heavy decode accelerators opened an Ultra tier at 1,000 TPS per user. Vera-Rubin NVL72 plus Groq delivered 35X the Premium-tier performance of Grace-Blackwell alone.

Why didn't AMD follow Nvidia into Groq?

AMD's MI accelerators share the decode-stage bottleneck that pushed Nvidia to disaggregate inference. The three remaining accelerator alternatives were unavailable for full acquisition:

  • Cerebras Systems: IPO'd earlier this year. Market cap $50.9 billion. Under $200 million quarterly revenue. A takeover would cost roughly $60 billion, or 300X annual revenue.
  • Graphcore: Acquired by SoftBank, owner of Arm.
  • SambaNova: Partnering with Intel.

That left AMD partnering with Cerebras for disaggregated inference rather than acquiring it. Cerebras plans to announce its fourth-generation WSE-4 wafer-scale engine later this year, paired with AMD's "Helios" rack systems built on "Verano" Epyc CPUs and "Altair" MI455X GPUs.

What makes Taalas different?

Taalas exited stealth in February, founded by engineers who previously worked at Tenstorrent. The company hardcodes AI model weights into ROM circuits on the chip, linked to on-chip SRAM blocks acting as a KV cache.

Taalas hardware specifications reported at launch:

  • HC1 generation: models up to 8 billion parameters per device
  • HC2 generation: models up to 20 billion parameters
  • Tens of chips linked together: models up to one trillion parameters

Initial benchmarks show much lower cost per token and lower latency versus Nvidia Blackwell B200 GPUs. Each model requires a new HC chiplet variant because two metal layers encoding the weights must change. Taalas says customizing an HC variant costs about 1/100th the price of training a new generative AI model.

What does AMD plan next?

An AMD statement said: "AMD plans to integrate the technology into its accelerator roadmap and develop system-level solutions with AMD Instinct GPUs."

AMD disclosed no timeline, packaging choice or integration cost.

With Cerebras out of reach for outright purchase, AMD paired a partnership it cannot buy with an outright acquisition it can. The four major model-specific accelerator vendors, Cerebras, Graphcore, SambaNova and Groq, now sit across four different corporate parents, with only Taalas available for a clean AMD takeover.

via The Next Platform (Source)

Filed under

  • amd
  • taalas
  • ai-inference
  • instinct
  • cerebras
Share this article:

More from Amara Osei

Amara Osei

Show full bio

Staff writer covering business strategy at Die Signal.

254 articles

Same lot · LOT-C1C6

« Previous articleNext article »