Switch language한국어
Back to the list

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OmegaUse-OfficeVal, a 100-task benchmark for long-horizon office-suite workflows.

  2. It uses privacy-preserved practitioner requests plus human labor-time and price-proxy signals to ground cost evaluation.

  3. Code-based verifiers were built from detailed rubrics to score task outputs more reliably.

  4. Frontier LLMs were cheaper and faster than human workers, but still fell short of human-level output quality.

Read the original