Switch language한국어
Back to the list

Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

TL;DR AI

Key summary

2 min read
  1. Video2GUI mines more than 500 million YouTube videos to extract grounded interaction trajectories and build WildGUI, a 12.7 million-trajectory dataset for GUI agent pretraining.

  2. Pretraining vision-language models on WildGUI delivered 5–20% gains, suggesting large-scale video mining can reduce reliance on costly manual annotation.

  3. The approach improved performance across web, mobile, and desktop benchmarks, including ScreenSpot-Pro, OSWorld-G, AndroidControl, CAGUI, OSWorld, and AndroidWorld.

  4. Models such as Qwen2.5-VL and Mimo-VL benefited from the dataset, highlighting its value for general-purpose GUI agents.

Read the original