RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

TL;DR AI
2 min readKey summary
Researchers introduced RSIBench-Data, a benchmark for testing LLM agents as data-centric researchers in a fixed post-training pipeline.
Across six benchmarks, four frontier agents showed occasional gains from feedback, but progress was often inconsistent and later revisions frequently failed to beat earlier best results.
The study found stronger runs were linked to accurate hypotheses, validation-based supervision, behavior-aligned data, and preserving strong checkpoints.
RSIBench-Data offers a way to measure whether AI systems can learn from failed runs to improve data choices, a key step toward recursive self-improvement.
