Understanding the Impact of Data Temporality on Large Language Model Pre-training
TL;DR AI
2 min readKey summary
A study found that LLMs pretrained on temporally ordered Common Crawl data were more current and time-aware than models trained on shuffled data.
Using a 7,000+ question benchmark, the ordered models showed better temporal grounding and factual freshness.
They stayed competitive on general language tasks, suggesting time order can improve reliability without sacrificing broad performance.
