LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
TL;DR AI
2 min readKey summary
Researchers introduced the Shannon Scaling Law, reframing LLM training as communication over a noisy channel.
On Pythia and OLMo2, it better captured non-monotonic performance drops from perturbations such as noise and quantization.
The model outperformed traditional power-law scaling laws in prediction and extrapolation.
The framework may help identify when more data or larger models stop improving results and start hurting them.
