Language-Switching Triggers Take a Latent Detour Through Language Models
TL;DR AI
2 min readKey summary
Researchers traced a Latin trigger in an 8B autoregressive language model that flips English responses into French.
The paper finds a three-stage circuit: early attention heads assemble the trigger, a latent subspace carries it, and the final-layer MLP converts it into French output.
This shows how a hidden backdoor can work below the surface, bypassing simple input checks and evading common defenses.
The finding highlights a new model-security risk in internal representations and language-switching attacks.
