Switch language한국어
Back to the list

New benchmark shows Claude Mythos and GPT-5.5 can develop real browser exploits autonomously

TL;DR AI

Key summary

2 min read
  1. Carnegie Mellon researchers released ExploitBench, a benchmark for real Google V8 vulnerabilities that measures progress toward arbitrary code execution.

  2. In tests, Claude Mythos Preview far outperformed GPT-5.5 in both assisted and fully autonomous exploit development, reaching the top tier on 21 of 41 bugs versus two.

  3. The stronger performance came at a much higher cost, underscoring a major capability-versus-efficiency gap across frontier models.

  4. The benchmark uses publicly known V8 bugs, so it does not yet measure discovery of new flaws or full weaponization.

  5. The results highlight growing concern that advanced AI models can already do substantial offensive security work.

Read the original