RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data

TL;DR AI
2 min readKey summary
Researchers used ATLAS to trace 1.45 million RLVR samples back to just 20 atomic sources, revealing heavy reuse across datasets and clear contamination risks.
Using source-level counterfactual attribution, they built DAPO++, a decontaminated training set designed to remove overlap and improve data provenance.
They also introduced a dataset quality score that aligns with downstream RLVR performance, helping predict which datasets are likely to train better models.
The findings show that RLVR benchmarks often share the same upstream data, making lineage tracking and contamination control essential for reliable training.
