OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
TL;DR AI
2 min readKey summary
Researchers introduced OmegaUse-OfficeVal, a 100-task benchmark for long-horizon office-suite workflows.
It uses privacy-preserved practitioner requests plus human labor-time and price-proxy signals to ground cost evaluation.
Code-based verifiers were built from detailed rubrics to score task outputs more reliably.
Frontier LLMs were cheaper and faster than human workers, but still fell short of human-level output quality.
