Switch language한국어
Back to the list

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OmegaUse-OfficeVal, a 100-task benchmark for office-suite workflows built from real practitioner requests with privacy-preserving adaptation.

  2. Each task includes human labor-time estimates, a price proxy, and code-based verifiers, making evaluation more realistic than simple task-success checks.

  3. Tests on frontier LLMs and a human baseline show models are faster and cheaper than people, but still lag behind human-level deliverable quality.

  4. The benchmark aims to measure whether LLM agents can handle long-horizon office work with human-like quality while accounting for time and cost.

Read the original