Learning When to Translate for Multilingual Reasoning

TL;DR AI
2 min readKey summary
Researchers introduced Luar, a reinforcement learning method that teaches reasoning language models when to translate non-English inputs into English.
Instead of translating everything, the model learns a selective policy: answer directly when comprehension is reliable, and translate only when needed.
The approach outperforms baseline training methods on multilingual benchmarks, with especially strong gains for low-resource languages.
The work suggests a more efficient path to multilingual reasoning by improving accuracy while avoiding unnecessary translation.
