Switch language한국어
Back to the list

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

TL;DR AI

Key summary

2 min read
  1. Researchers introduced LongTraceRL, a reinforcement-learning framework for long-context reasoning in LLMs.

  2. It creates harder training examples from search-agent trajectories and knowledge-graph multi-hop questions, adding tiered distractors.

  3. The method uses entity-level rubric rewards only for correct answers, encouraging more evidence-grounded reasoning.

  4. Across five benchmarks and 4B–30B models, LongTraceRL beat strong baselines on long-context tasks.

Read the original