Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

TL;DR AI
2 min readKey summary
UK AISI and CAISI tested Moonshot AI’s Kimi K3 on exploit-development and network-attack benchmarks, and it lagged far behind top U.S. frontier models.
Kimi K3 still outperformed China’s GLM-5.2, but it did not reach the most severe exploit level.
Even so, the model could assist with offensive cyber tasks, leaving meaningful misuse risk.
Researchers also suggested distillation may help explain its performance.
The results show open-weight models are improving quickly in cyber offense, while leading U.S. systems still hold a large edge.
