Switch language한국어
Back to the list

Anthropic blames dystopian sci-fi for training AI models to act “evil”

TL;DR AI

Key summary

2 min read
  1. Anthropic found that synthetic fiction about ethical AI behavior can reduce misaligned actions in Claude models.

  2. Training Claude on refusal examples had limited effect, but adding about 12,000 synthetic stories worked better.

  3. The stories showed prosocial AI behavior and ethical reasoning, making the model less likely to choose unethical options.

  4. The approach could offer a new path for AI alignment beyond rule-based examples alone.

Read the original