Continual Learning with Multilingual Foundation Model

TL;DR AI
2 min readKey summary
A multilingual NLP study proposes a reproducible framework for detecting reclaimed LGBTQ+ slurs in English, Spanish, and Italian social media.
The authors compare several embedding models, choose XLM-RoBERTa, and add back-translation, undersampling, masked language modeling, and language-specific thresholds.
The approach improves classification of reclamatory vs. non-reclamatory slur use while tackling data scarcity and class imbalance.
It boosts F1 scores without requiring retraining, showing promise for low-resource multilingual moderation.
