AMD’s vLLM-ATOM Plugin Supercharges DeepSeek-R1, Kimi-K2, and gpt-oss-120B AI LLM Inference on Instinct MI350 and MI400 Accelerators

TL;DR AI
2 min readKey summary
AMD launched vLLM-ATOM, a plugin backend for vLLM that adds AMD-optimized model code and GPU kernels for Instinct MI350 and MI400 accelerators.
The plugin supports multiple LLM and VLM architectures, including DeepSeek-R1, Kimi-K2, and gpt-oss-120B, while preserving existing vLLM APIs.
AMD says the approach will speed up production inference and serving without forcing workflow changes for current vLLM users.
The company also plans to upstream the improvements later, helping spread the optimizations across the open-source ecosystem.

