AMD, Cerebras partner on AI inference system targeting 5x efficiency gain

Cerebras shares rose about 4% after the company said its chips will be used in AMD Helios systems in Cerebras data centers later this year, extending a broader push into ultra-low-latency AI infrastructure.

Summary

AMD and Cerebras Systems unveiled a technical partnership to build a disaggregated AI inference platform that combines AMD Helios rackscale systems with the Cerebras Wafer-Scale Engine. The companies said the setup is designed for ultra-low-latency AI inference while increasing throughput and efficiency, with a claimed 5x gain in tokens per second per watt. Cerebras CEO Andrew Feldman said at AMD's AI conference in San Francisco that Cerebras chips will be used in AMD's Helios AI systems installed in Cerebras data centers starting later this year, and that server buyers will also be able to configure AMD systems with Cerebras' wafer-scale chips. Cerebras shares rose about 4% on Thursday to trade at $219.80 after the announcement. The partnership underscores how AI infrastructure providers are trying to balance response speed, power efficiency and flexibility as demand grows for real-time applications. It also comes after Cerebras announced in January a deal with OpenAI worth over $10 billion to deliver 750 megawatts of computing power through 2028.

Terms & Concepts
  • AI inference: The stage when a trained AI model processes live requests and generates outputs.
  • ultra-low latency: Very fast response time, important for AI applications that need answers almost immediately.
  • wafer-scale chips: Processors built using an entire silicon wafer to deliver very large amounts of compute performance.