Confidence-Adaptive SwiGLU for Mixture-of-Experts
TL;DR AI
2 min readKey summary
Researchers introduced κ-SwiGLU, a confidence-aware SwiGLU variant for Mixture-of-Experts Transformers.
It adapts gate sharpness using token-level routing confidence from the router logit.
On FineWeb-Edu across 8- to 28-layer MoE Transformers, it improved mean CORE performance.
The gains came with negligible extra parameters and only a small compute overhead.
