SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

TL;DR AI
2 min readKey summary
Researchers introduced SaaS-Bench, a benchmark covering 106 real tasks across 23 SaaS products and six professional domains.
It tests computer-use and LLM agents on realistic workplace workflows, including long-horizon and cross-application actions.
Top agents complete fewer than 4% of tasks end to end, revealing major weaknesses in planning, state tracking, and recovery.
The benchmark offers a more realistic measure of AI agents in enterprise software and underscores why current systems still fall short.
