Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

TL;DR AI
2 min readKey summary
Researchers introduced CUActSpot, a benchmark for computer-use agents that covers GUI, text, table, canvas, and natural-image interactions.
The paper argues current agents fail on sparse, long-tail interaction types, especially complex GUI actions such as drag-and-draw workflows.
It also proposes a renderer-based synthesis pipeline to generate scenes, instructions, and action traces for training and evaluation.
A model trained on this corpus, Phi-Ground-Any-4B, reportedly beats open-source models under 32B parameters.
The authors plan to release the benchmark, synthetic data, code, and models to help improve real-world automation reliability.
