AMD agrees to acquire inference chip startup Taalas

AMD agrees to acquire inference chip startup Taalas

The deal adds model-specific inference silicon to AMD's AI stack as it pushes further into a fast-growing market for running trained models, where Nvidia has also been expanding its specialization strategy.

Fact Check
AMD's official investor-relations press release directly confirms the definitive agreement to acquire Taalas, a startup specializing in AI inference silicon, aligning with the claim that the deal adds specialized inference silicon to AMD's accelerator roadmap. Reuters and CNBC independently corroborate the acquisition, its inference focus, and integration into AMD's Instinct/Helios roadmap. Reuters also confirms the broader AI-stack deal context including MK1, MEXT, and FastFlowLM.
    Reference123
Summary

Advanced Micro Devices said Thursday it will acquire Toronto startup Taalas for an undisclosed amount, adding a model-specific AI inference technology designed to bypass the memory-movement bottleneck that limits token generation on conventional GPUs. Taalas hard-wires a model's weights directly into silicon, so its first chip, HC1, runs Meta's Llama 3.1 8B and no other base model. The company says that design sharply increases speed by avoiding repeated transfers of model weights between memory and compute. Taalas has said HC1 generates roughly 17,000 tokens per second per user, while EE Times reported seeing more than 15,000 in a public demo. Taalas' own testing compared that with about 350 on Nvidia's Blackwell hardware. The startup also says its system costs 20 times less to build and uses 10 times less power because it does not require HBM, advanced packaging or liquid cooling, though those figures remain company estimates. The trade-off is flexibility: switching to a different model requires making a new chip. Taalas says most of the design can remain unchanged and only a small portion must be modified to encode a new model, with TSMC able to manufacture an updated version in about two months. CEO Ljubisa Bajic told EE Times the company made "painful tradeoffs in flexibility for the sake of economics and speed." AMD has framed inference as a major growth opportunity, projecting the market for running trained models may grow more than 80% annually. Taalas is AMD's fourth inference deal since November, after inference software firm MK1 and two smaller startups, and the technology is expected to feed into Helios, AMD's rack-scale AI system. Microsoft last month agreed to deploy Helios on Azure to run frontier-model inference. The acquisition also underscores intensifying competition with Nvidia, which built its lead in AI training and has been making its own moves in specialized inference hardware.

Terms & Concepts
  • inference: The stage of AI deployment where a trained model is used to generate responses or predictions for users.
  • HBM: High-bandwidth memory, a fast memory technology used in AI chips to move large amounts of data quickly.
  • rack-scale: A system design that combines many computing components across an entire server rack rather than relying on a single chip or server.