WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

TL;DR AI
2 min readKey summary
Researchers introduced WorkSurface-Bench, a 1,151-task benchmark for enterprise agents to route questions across documents, tables, graphs, and cross-surface sources.
Across four model backbones and six agent setups, routing was often near-perfect with gold tool access, but answer accuracy lagged well behind.
The result shows that choosing the right knowledge surface is necessary but not sufficient for correct enterprise-agent answers.
The team also released the dataset, scoring code, and agent harness to support further evaluation.
