GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
TL;DR AI
2 min readKey summary
GoLongRL is a fully open-source long-context RLVR framework with a 23K-sample dataset spanning nine capability types.
It pairs the data with TMN-Reweight, a method that stabilizes heterogeneous multitask rewards during training.
Reported results beat a closed-source dataset baseline and approach the performance of much larger leading models.
The work offers public data, code, and a training recipe for stronger long-context reasoning and broader task coverage.
