Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers
TL;DR AI
2 min readKey summary
The study introduces a human-annotated code-switching retrieval benchmark plus an 11-task evaluation suite.
Statistical, dense, and late-interaction retrievers all show major performance drops on mixed-language queries.
The paper finds an embedding-space mismatch between pure-language and code-switched text.
Standard multilingual fixes, including vocabulary expansion, provide only partial improvement, leaving a key search gap unresolved.
