LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

TL;DR AI
2 min readKey summary
Researchers introduced LongTraceRL, a reinforcement-learning method for long-context reasoning in large language models.
It creates harder training contexts from search-agent trajectories and tiered distractors, then applies entity-level rubric rewards on correct answers.
Across multiple model sizes and five benchmarks, LongTraceRL outperformed strong baseline methods.
The approach is designed to help models better locate and use evidence in large documents and multi-hop tasks.
