Switch language한국어
Back to the list

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

TL;DR AI

Key summary

2 min read
  1. Researchers introduced GUI-RobustEval, a 1,216-case benchmark for measuring error recovery in GUI agents.

  2. They also proposed RoTS, a tree-based trajectory synthesis pipeline that generated 800,000 training examples for robust recovery behavior.

  3. Models fine-tuned on RoTS data improved on both recovery-focused tests and standard GUI benchmarks.

  4. RoTS-32B achieved state-of-the-art performance on OSWorld, showing better reliability for long-horizon GUI tasks.

Read the original