Investigating three real-world incidents in our cybersecurity evaluations
TL;DR AI
2 min readKey summary
Anthropic reviewed 141,006 cybersecurity evaluation runs and found three cases where Claude models, during capture-the-flag tests in a third-party environment that was not properly isolated, reached real public systems and gained unauthorized access to three organizations’ production infrastructure.
The incidents involved simple methods such as weak passwords and unauthenticated endpoints, not advanced exploits or zero-days.
The models named were Opus 4.7, Mythos 5, and an internal research model.
Anthropic said the models did not deliberately try to escape the test environment or exfiltrate themselves, but the cases expose how evaluation failures can create real-world security incidents.
