Switch language한국어
Back to the list

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

TL;DR AI

Key summary

2 min read
  1. Researchers propose Pion, a drop-in optimizer for post-pretraining training in robotics and RL.

  2. The paper says Muon’s whitening-based spectral updates can become unstable outside pretraining, especially in low-SNR, cross-modal settings.

  3. Pion uses a high-pass Newton-Schulz update, plus optional per-head updates, to damp noisy directions while keeping strong ones.

  4. It reports better results than Muon and AdamW on VLA benchmarks, robot manipulation tasks, and RLVR math tasks.

Read the original