Scaling Native Multimodal Pre-Training From Scratch
TL;DR AI
2 min readKey summary
Researchers mapped how to scale vision-language models trained from scratch under a fixed compute budget.
They found predictable compute-loss curves and power-law rules for choosing optimal model size and token count.
Language and multimodal objectives scale differently, and the data mix strongly affects multimodal efficiency.
Native multimodal pre-training can also improve spatial reasoning and multimodal in-context learning.
