AMD partners with Cerebras to unveil Helios, a new system aimed at improving AI inference and challenging Nvidia’s dominance.
AMD partners with Cerebras to unveil Helios, a new system aimed at improving AI inference and challenging Nvidia’s dominance.
AMD is making a significant move in the realm of artificial intelligence, asserting that reliance on a single chip technology is not the future. During an announcement on Thursday, CEO Lisa Su confirmed the company’s collaboration with Cerebras, a chip startup focused on redefining AI inference—the method through which AI models generate responses.
As the demand for sophisticated chip designs surges, especially in light of the AI boom, AMD’s partnership with Cerebras comes as a strategic response. The contemporary trend among chipmakers, often referred to as “disaggregated inference,” allows for workload distribution across various hardware types, diverging from traditional methods that utilize a single type of hardware for both processing requests and producing answers. AMD contends that these two tasks are fundamentally different, necessitating a specialized approach.
Helios, AMD’s newest server system, is engineered to handle extensive volumes of requests efficiently, while Cerebras contributes its massive, wafer-sized chip renowned for delivering near-instantaneous responses. This partnership will see Helios deployed within Cerebras’ data centers before the end of the year.
The surge in demand for advanced chips has prompted companies like AMD, NVIDIA, and Broadcom to ramp up production to meet the escalating needs of the market. AMD is stepping into this competitive landscape, particularly as companies shift their focus from merely training AI models to actual deployment in practical environments.
The alignments with Cerebras reflect a broader trend noted by analysts, who emphasize that the shortcomings of existing chip architectures are catalyzing the transition towards disaggregated inference. A report from UBS as of June indicated that competitors, including NVIDIA—with their acquisition of AI hardware startup Groq—and Amazon Web Services, are also adopting similar setups to enhance performance while reducing costs.
Despite the exciting prospects, UBS also pointed out the challenges associated with disaggregated inference, particularly concerning the orchestration of various chips to function cohesively within a system.
At the recent Advancing AI event, AMD introduced Helios, which bundles diverse types of AI chips to compete directly with NVIDIA’s Vera Rubin NVL72 rack. Among those using AMD’s infrastructure are major players such as OpenAI, Meta, Microsoft, Oracle, and Anthropic, with whom AMD recently established a multi-billion dollar partnership.
AMD is not shying away from taking shots at its competition. During the same event, the company claimed that Helios offers up to 30% more inference tokens per dollar compared to NVIDIA’s Vera Rubin NVL72 rack, asserting, “Every Helios can deliver more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks,” according to Su.
Keep in touch with our news & offers
Subscribe to Our Newsletter
Thank you for subscribing to the newsletter.
Oops. Something went wrong. Please try again later.












