CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations
TL;DR AI
2 min readKey summary
Researchers extended contrastive preference tuning from English to 14 languages for multilingual LMs.
Using self-generated outputs and an English-trained reward model, the method improved many structured and open-ended tasks.
The approach worked best with on-policy data and helped avoid catastrophic forgetting seen with supervised fine-tuning.
It offers a practical path to multilingual alignment without costly language-specific preference annotations.
