Switch language한국어
Back to the list

Understanding the Impact of Data Temporality on Large Language Model Pre-training

TL;DR AI

Key summary

2 min read
  1. A study found that LLMs pretrained on temporally ordered Common Crawl data were more current and time-aware than models trained on shuffled data.

  2. Using a 7,000+ question benchmark, the ordered models showed better temporal grounding and factual freshness.

  3. They stayed competitive on general language tasks, suggesting time order can improve reliability without sacrificing broad performance.

Read the original