MobileMoE: Scaling On-Device Mixture of Experts
TL;DR AI
2 min readKey summary
Researchers introduced MobileMoE, a family of compact on-device mixture-of-experts language models for smartphones.
Built with a mobile-aware scaling law and multi-stage training recipe, the models match or beat strong dense and MoE baselines.
MobileMoE delivers better benchmark performance with fewer inference FLOPs and faster smartphone execution, including efficient INT4 deployment.
The results suggest sparse expert models can be practical for mobile AI, opening a new efficiency-performance frontier.
