AMD and Cerebras partner on low-latency, high-throughput AI inference — EPYC processors in Helios rack-scale infrastructure paired with Cerebras' Wafer-Scale Engine (WSE) solutions

TL;DR AI
2 min readKey summary
AMD and Cerebras are co-developing a disaggregated AI inference platform that splits work across specialized hardware.
AMD’s EPYC-based Helios racks and Instinct MI400 GPUs will handle prompt processing and context-heavy prefill, while Cerebras Wafer-Scale Engine systems will generate latency-sensitive tokens.
The service is planned to launch on Cerebras Cloud in the second half of 2026.
The approach aims to improve efficiency, throughput, and cost, and could pressure Nvidia’s inference strategy.



