Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
TL;DR AI
2 min readKey summary
Video2GUI mines more than 500 million YouTube videos to extract grounded interaction trajectories and build WildGUI, a 12.7 million-trajectory dataset for GUI agent pretraining.
Pretraining vision-language models on WildGUI delivered 5–20% gains, suggesting large-scale video mining can reduce reliance on costly manual annotation.
The approach improved performance across web, mobile, and desktop benchmarks, including ScreenSpot-Pro, OSWorld-G, AndroidControl, CAGUI, OSWorld, and AndroidWorld.
Models such as Qwen2.5-VL and Mimo-VL benefited from the dataset, highlighting its value for general-purpose GUI agents.
