Nvidia paid $20 billion for SRAM decoding – AMD just partnered for it instead


  • AMD and Cerebras will share inference across two machines, with Helios racks handling fast processing and Wafer-Scale Engine generating tokens, available through Cerebras Cloud in H2 2026
  • Nvidia is also doing something similar by licensing AI chip startup Groq’s SRAM decoding technology for $20 billion
  • The move means that AMD and Cerebras claim 5x higher tokens per watts compared to a standalone Cerebras WSE configuration

AMD and Cerebras Systems have announced a technical partnership that pairs the former’s Helios rackscale system with the latter’s Wafer-Scale Engine in what both companies call a disaggregated inference solution.

The move has enabled a combined AMD Helios and Cerebra’s WSE configuration to deliver up to five times as many tokens per second per watt (TPS/W) in internal testing by both chip designers.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top