Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
TL;DR AI
2 min readKey summary
Researchers introduced GUI-RobustEval, a 1,216-case benchmark for measuring error recovery in GUI agents.
They also proposed RoTS, a tree-based trajectory synthesis pipeline that generated 800,000 training examples for robust recovery behavior.
Models fine-tuned on RoTS data improved on both recovery-focused tests and standard GUI benchmarks.
RoTS-32B achieved state-of-the-art performance on OSWorld, showing better reliability for long-horizon GUI tasks.
