Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
TL;DR AI
2 min readKey summary
Researchers propose Pion, a drop-in optimizer for post-pretraining training in robotics and RL.
The paper says Muon’s whitening-based spectral updates can become unstable outside pretraining, especially in low-SNR, cross-modal settings.
Pion uses a high-pass Newton-Schulz update, plus optional per-head updates, to damp noisy directions while keeping strong ones.
It reports better results than Muon and AdamW on VLA benchmarks, robot manipulation tasks, and RLVR math tasks.
