Second-Order Injection: Attacking the Evaluator in LLM Safety Monitors

TL;DR AI
2 min readKey summary
Researchers found that attacker-controlled session text can hijack LLM safety evaluators, making harmful content appear safe.
Across qwen2.5:3b, mistral, and phi3:mini, tuned injection vectors fully overrode verdicts and transferred across model families.
Even dual-evaluator setups lost their divergence under symmetric injection, weakening the intended defense.
Prompt-level sanitization helped only partially and did not solve the core issue that the evaluator itself can be manipulated.

