Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
TL;DR AI
2 min readKey summary
A new study says larger models learn more because they face less task interference and can keep rare-task features better.
In synthetic multi-task experiments and OLMo models from 4M to 4B parameters, bigger models handled infrequent and complex tasks more effectively.
The researchers found that larger models can separate tasks more cleanly, reducing gradient interference and improving task retention.
They also seem able to allocate more capacity to common tasks without losing harder ones, which smaller models struggle to do.
The results help explain why scale improves performance on rare and complex behaviors and can inform model-size and data-mix choices.
