AMD introduced the Helios AI rack-scale system designed to support large AI labs, with deployments scheduled for later in 2026 [1, 2, 3]. The system targets AI inference workloads by combining AMD chips optimized for request processing and Cerebras wafer-scale chips built for ultra-low latency AI responses [1, 3].
The partnership between AMD and chip startup Cerebras, announced July 23 at AMD’s Advancing AI conference in San Francisco, entails integrating Cerebras chips into Helios systems installed directly in Cerebras data centers starting later this year [1, 2, 3]. Cerebras’ CEO Andrew Feldman said, "When something's a necessity, people want to use it, and they want to use it quickly" [1].
Helios counts customers such as OpenAI, Microsoft, Meta, Oracle, and Anthropic, reflecting strong demand among leading AI firms [2, 3]. AMD claimed Helios delivers up to 30% more inference tokens per dollar compared to Nvidia’s Vera Rubin NVL72 rack system [3]. Lisa Su, AMD Chair and CEO, stated, "Every Helios can deliver more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks" and described Helios as "the tech industry’s highest performance AI rack, built to train and run the most demanding frontier models in the world at massive scale" [2, 3].
In May 2026, Cerebras went public with volatile stock trading between $161 and $386.34; after the AMD partnership news, shares rose about 4% to $219.80 [1]. AMD also announced the Venice-X CPU for data centers, planned for launch in 2027 [2].
Separately, AMD and Anthropic announced a strategic partnership to deploy up to two gigawatts of GPUs via the Helios system, indicating further expansion for the AI infrastructure platform [2].
Deployment of Helios systems, featuring integrated Cerebras wafer-scale chips, will begin later this year at Cerebras data centers, marking a significant step in AMD’s AI infrastructure push [1, 3].