Switch language한국어
Back to the list

LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

TL;DR AI

Key summary

2 min read
  1. Researchers introduced the Shannon Scaling Law, reframing LLM training as communication over a noisy channel.

  2. On Pythia and OLMo2, it better captured non-monotonic performance drops from perturbations such as noise and quantization.

  3. The model outperformed traditional power-law scaling laws in prediction and extrapolation.

  4. The framework may help identify when more data or larger models stop improving results and start hurting them.

Read the original