ORCA-bench: How Ready Are Language Model Agents for Oncall?

TL;DR AI
2 min readKey summary
Researchers introduced ORCA-bench, a public benchmark for coding agents on production-style oncall root cause analysis.
It uses real telemetry data, source code, and 1,079 incident tasks to test how agents handle realistic incidents.
Across five frontier agents, accuracy was low and hallucinated root causes were common.
Performance dropped further without source-code access, suggesting current agents are not yet reliable for production reliability work.
