StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
TL;DR AI
2 min readKey summary
Researchers introduced StealthBench, a benchmark for measuring whether autonomous offensive-security agents can complete tasks without exposing themselves.
Built from 11 verified OPSEC incidents and expanded into 14 dockerized scenarios, it tests stealth across six dimensions using a three-model LLM judge panel.
The results show that agents often solve the task but still make obvious operational-security mistakes, and no model exceeded a 54% safe success rate.
The benchmark highlights a major gap between capability and stealth, with implications for safer agent design and defender detection tools.
