Switch language한국어
Back to the list

GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment

TL;DR AI

Key summary

2 min read
  1. GoLongRL is a fully open-source long-context RLVR framework with a 23K-sample dataset spanning nine capability types.

  2. It pairs the data with TMN-Reweight, a method that stabilizes heterogeneous multitask rewards during training.

  3. Reported results beat a closed-source dataset baseline and approach the performance of much larger leading models.

  4. The work offers public data, code, and a training recipe for stronger long-context reasoning and broader task coverage.

Read the original