OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

TL;DR AI
2 min readKey summary
Researchers introduced OmegaUse-OfficeVal, a 100-task benchmark for office-suite workflows built from real practitioner requests with privacy-preserving adaptation.
Each task includes human labor-time estimates, a price proxy, and code-based verifiers, making evaluation more realistic than simple task-success checks.
Tests on frontier LLMs and a human baseline show models are faster and cheaper than people, but still lag behind human-level deliverable quality.
The benchmark aims to measure whether LLM agents can handle long-horizon office work with human-like quality while accounting for time and cost.
