SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarking
TL;DR AI
2 min readKey summary
Researchers introduced SimuWoB, a fully synthetic benchmark for mobile GUI agents with 120 app-like tasks.
The benchmark uses automatically generated rewards and URL-accessible, backend-free virtual environments for reproducible evaluation.
Testing showed state-of-the-art agents still achieve low success rates, especially on long-horizon, multi-step tasks.
The results highlight a more realistic way to measure mobile agent performance and expose current limitations on complex workflows.
