- AMD and Cerebras will share inference across two machines, with Helios racks handling fast processing and Wafer-Scale Engine generating tokens, available through Cerebras Cloud in H2 2026
- Nvidia is also doing something similar by licensing AI chip startup Groq’s SRAM decoding technology for $20 billion
- The move means that AMD and Cerebras claim 5x higher tokens per watts compared to a standalone Cerebras WSE configuration
AMD and Cerebras Systems have announced a technical partnership that pairs the former’s Helios rackscale system with the latter’s Wafer-Scale Engine in what both companies call a disaggregated inference solution.
The move has enabled a combined AMD Helios and Cerebra’s WSE configuration to deliver up to five times as many tokens per second per watt (TPS/W) in internal testing by both chip designers.
The move aims to address a Cerebra WSE efficiency challenge by porting fast processing to AMD’s rackscale offering.
Latest videos fromTechRadar
An efficiency gain-centric game?
Both AMD and Cerebras Systems are painting the news as a win, and it may well be, given the latter’s in-game efficiency gains and the former’s ability to access SRAM decoding technology without spending the $20 billion Nvidia shelled out late last year for a non-exclusive deal.
However, it should be noted that the efficiency requirements of 5 tokens per second per watts compared to an existing Cerebra’s WSE (Wafer-Scale Engine) as a baseline while running the open-source Kimi 2.6 1T model, which makes them impressive, but without a direct comparison to figures for an Nvidia rack-scale that lacks efficiency, especially when efficiency is the metric context.
The idea itself is sound and well-established in the industry, with WSE known to struggle with the ‘prefill’ part of the equation, while handling the ‘decode’ segment relatively well, essentially replacing AMD’s hardware where Cerebras’ equipment falls short.
However, the choice of Kimi 2.6 deserves a second look. Released on April 20, 2026, Moonshot AI’s model is a mix of expert designs with a trillion total parameters but only 32 billion actives per token, and it ships natively in INT4. At INT4, the full weight set runs to around 500GB. A single Cerebras wafer holds 44 GB. Even before KV cache, a Cerebras-only implementation needs somewhere north of a dozen wafers just to hold the model, while a Helios rack can hold about sixty times that.
That asymmetry means that the five-fold figure is measured on a model that is close to the least favorable for a WSE-only configuration. A dense model small enough to sit on a handful of wafers could flatter the Cerebras considerably more. None of this makes the number wrong, but it does warrant further testing to demonstrate both its strengths and weaknesses for different models.
A partnership without numbers, for now
More importantly, the absence of any financial information may well be a future story, especially at a time when there are growing concerns about ‘circular finance’ in an industry where Nvidia’s recent move to stop OpenAI’s data center purchases was seen as a net negative by Wall St, which is already worried about AI spending and the sustainability of such transactions.
AMD has also in the past (and more recently with Anthropic) linked purchases of its own hardware to investments or equity stakes it would take in AI companies, moves the market welcomed in the past but may view with a little more hostility lately.
The announcement comes at a time when Cerebras may need it more than AMD: Listed on the Nasdaq in May at $185, Cerebras opened at $350 and closed its first day at $311.07 before falling back to around $227 in late June 2026.
AMD stock, on the other hand, is up 121.48% year-to-date (YTD) as investors continue to bet heavily on its new Instinct AI processors and the Cerebras partnership allows it to further consolidate its gains as this could be seen as another vote of confidence in its current direction by one of its potential customers.
Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews and opinions in your feeds.



