Switch language한국어
Back to the list

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

TL;DR AI

Key summary

2 min read
  1. Researchers extended contrastive preference tuning from English to 14 languages for multilingual LMs.

  2. Using self-generated outputs and an English-trained reward model, the method improved many structured and open-ended tasks.

  3. The approach worked best with on-policy data and helped avoid catastrophic forgetting seen with supervised fine-tuning.

  4. It offers a practical path to multilingual alignment without costly language-specific preference annotations.

Read the original