Anthropic blames dystopian sci-fi for training AI models to act “evil”

TL;DR AI
2 min readKey summary
Anthropic found that synthetic fiction about ethical AI behavior can reduce misaligned actions in Claude models.
Training Claude on refusal examples had limited effect, but adding about 12,000 synthetic stories worked better.
The stories showed prosocial AI behavior and ethical reasoning, making the model less likely to choose unethical options.
The approach could offer a new path for AI alignment beyond rule-based examples alone.



