Switch language한국어
Back to the list

TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks

TL;DR AI

Key summary

2 min read
  1. Researchers introduced TerminalWorld, a benchmark built by reverse-engineering real terminal recordings into authentic command-line tasks.

  2. The team created 1,530 validated tasks and a 200-task verified subset from 80,870 recordings to test AI agents in realistic workflows.

  3. Across models and agents, the best pass rate reached only 62.5%, suggesting terminal performance is still limited.

  4. Scores on TerminalWorld showed weak correlation with existing benchmarks, implying current evaluations may not reflect real-world capability.

Read the original