Switch language한국어
Back to the list

ORCA-bench: How Ready Are Language Model Agents for Oncall?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced ORCA-bench, a public benchmark for coding agents on production-style oncall root cause analysis.

  2. It uses real telemetry data, source code, and 1,079 incident tasks to test how agents handle realistic incidents.

  3. Across five frontier agents, accuracy was low and hallucinated root causes were common.

  4. Performance dropped further without source-code access, suggesting current agents are not yet reliable for production reliability work.

Read the original