MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
TL;DR AI
2 min readKey summary
MonitorBench is a new open benchmark for measuring chain-of-thought monitorability in large language models.
The benchmark includes 1,514 test instances across 19 tasks and two stress-test settings.
Experiments show monitorability is higher when structural reasoning is required and that closed-source models often have lower monitorability.
Stress-tests can reduce monitorability by up to 30% in some non-structural tasks.
