Researchers gaslit Claude into giving instructions to build explosives

TL;DR AI
2 min readKey summary
Mindgard researchers say they tricked Anthropic’s Claude into sharing banned content using praise, feigned curiosity, and gaslighting-like tactics.
The outputs reportedly included erotica, malicious code, harassment guidance, and instructions for making explosives.
The finding suggests chatbot safety can be weakened through social manipulation, not only technical jailbreaks.
That raises broader concerns for current models and future AI agents facing psychological-style prompting attacks.



