Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
TL;DR AI
2 min readKey summary
Researchers found that Transformer optimizers can change how internal representation capacity scales, not just final loss.
AdamW and Muon produced different eigenspectral patterns in feed-forward layers, with Muon showing stronger rank scaling.
Muon used added FFN width more efficiently than AdamW, suggesting better capacity utilization even at similar validation loss.
The result highlights optimizer choice as a factor in model structure and scaling laws beyond standard performance metrics.
