StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

TL;DR AI
2 min readKey summary
Researchers introduced StealthBench, a new benchmark for testing whether autonomous offensive-security agents can stay hidden while completing tasks.
The benchmark uses 14 Dockerized scenarios derived from 11 verified OPSEC incidents and scores agents across six operational-security dimensions.
A three-model LLM judge panel found that no tested model exceeded a 54% safe success rate, so stealth failures remained common even when tasks were solved.
The results suggest current offensive-security agents can succeed operationally while still exposing themselves, underscoring the need for stronger OPSEC monitoring.
