CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising

TL;DR AI
2 min readKey summary
Researchers introduced CDAE, a lightweight model that improves the robustness of BERT-based sentence embeddings.
CDAE combines contrastive and reconstruction objectives to reduce sensitivity to perturbations like synonym substitution, masking, and word dropout.
Across multiple perturbation strengths, CDAE preserved embedding similarity better than baseline BERT embeddings and SimCSE.
More stable sentence representations could make downstream NLP systems more reliable on slightly altered or noisy text.
