AMD Fires Back At NVIDIA’s Groq Bet, Fuses The Cerebras Wafer-Scale Engine With Helios For 5x Higher Tokens Per Second Per Watt

TL;DR AI
2 min readKey summary
AMD is partnering with Cerebras to combine Helios rack-scale systems with the Wafer-Scale Engine for AI inference.
The goal is to improve latency, efficiency, and token throughput, with claims of up to 5x better tokens per second per watt.
The move positions AMD as a stronger alternative to NVIDIA-linked inference stacks and a rival to Groq in AI acceleration.
If successful, the collaboration could give customers a more flexible high-performance platform for low-latency token generation.



