Lisa Su pushes disaggregated inference with Cerebras, aiming Helios at Nvidia’s Vera Rubin
AMD’s new Cerebras pact is a direct bet that AI inference should be split across specialized chips, not one dominant rack.

AMD President and CEO Lisa Su announced AMD’s teaming with chip startup Cerebras to pursue disaggregated inference for AI inference workloads. The move challenges Nvidia’s hardware dominance by positioning AMD’s Helios server system as a higher-performance, lower-cost alternative.
AMD President and CEO Lisa Su announced Thursday that AMD is teaming up with chip startup Cerebras on a new approach to AI inference, the process of generating responses from AI models. The bet is explicitly about structure, not just speed: AMD is pushing “disaggregated inference,” where different parts of the workload run on different types of hardware.
This matters because AMD is not just joining the AI infrastructure scramble. It is also taking a shot at Nvidia’s rack ecosystem with a specific performance and cost claim. At AMD’s Advancing AI event, the company said Helios delivers up to 30% more inference tokens per dollar than Nvidia’s Vera Rubin NVL72 rack. In other words, AMD wants customers to pay less for the same output, or get more output for the same budget, while also scaling across thousands of racks.
So what is “disaggregated inference,” and why is AMD treating it like a future-proofing strategy? Traditionally, one piece of hardware handled both prompt processing and answer generation. AMD argues those jobs are fundamentally different jobs. In its framing, Helios is designed to process huge volumes of requests, while Cerebras’ giant, wafer-sized chip specializes in generating near-instantaneous responses.
That split is more than an engineering detail. It is a procurement and architecture decision for any company building or serving large language models, because it changes what you buy and how you deploy. Instead of buying a single type of accelerator and hoping it is optimal for every step, disaggregation pushes the market toward mixing specialized chips that can complement each other. And with demand for chips from companies like AMD, Nvidia, and Broadcom skyrocketing in the AI boom, those differences quickly turn into real budgeting leverage.
AMD is framing this as part of a broader industry shift that has already started to form. Analysts cited in the original reporting point to UBS writing in June that limitations of current architectures are driving a shift toward disaggregated inference. UBS also wrote that Nvidia, through its integration of AI hardware startup Groq, and Amazon Web Services are pursuing similar setups to improve efficiency and lower costs. The shared theme: as the industry shifts focus from training models to putting them to work, inference becomes the dominant cost center, and architectures that waste cycles get squeezed.
Disaggregation also creates new headaches, and AMD can’t hand-wave that away. UBS pointed to “orchestration” as a major challenge, meaning the problem of getting different chips to work together seamlessly. This is where deals like AMD and Cerebras become less like a press release and more like a systems integration project. Timing is part of the story too: the partnership is set to bring Helios into Cerebras’ data centers later this year.
At the same time, AMD is treating Helios as its direct response to Nvidia’s server rack strategy. The original reporting notes that at Advancing AI, AMD unveiled Helios, which bundles several types of AI chips, described as the company’s answer to Nvidia’s Vera Rubin NVL72 rack. AMD also listed AI labs and cloud giants using AMD’s infrastructure, including OpenAI, Meta, Microsoft, Oracle, and Anthropic. AMD previously announced a multibillion-dollar infrastructure partnership with Anthropic on Wednesday, reinforcing that the infrastructure layer is where AMD wants to win mindshare and spend.
And then there is the strategic implication for everyone else building AI stacks. Nvidia dominates chip design for AI training, and the competition has intensified as AI companies move from training to serving. AMD’s approach suggests it thinks the “winning” inference infrastructure will not be monolithic. It will be composable, where different hardware types are allocated to different bottlenecks. For executives, that changes how you think about vendor risk, cost modeling, and roadmap alignment, because your next architecture might be defined as much by software orchestration and partner compatibility as by raw accelerator specs.
If AMD’s Helios-Cerebras play lands, it puts pressure on the entire inference supply chain: customers will ask whether their current rack-based approach is leaving tokens per dollar on the table. If it stumbles, the critique will be blunt, because disaggregated inference is supposed to solve efficiency and cost, and it can only do that if orchestration works at scale. Either way, AMD is making a clear bet that the next big shift in AI is not another incremental chip. It is a different way to assemble the hardware that delivers the output.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

Black Forest Labs launches FLUX 3 with 20-second audio-video, but no pricing yet
Enter gated early access for video and action, while image waits, benchmarks stay preliminary, and costs are a mystery.

AMD’s Helios claims 30% per-dollar lead over Nvidia’s Vera Rubin
The 72-GPU rack looks like a head-on spec war, but real-world FLOPS and power efficiency decide the winner.

Samsung commits to Wear OS updates for Galaxy Watch 9 and Ultra 2 until 2031
A five-year software promise reshapes wearables planning for buyers, app makers, and the companies betting on stickier ecosystems.

