Switch language한국어
Back to the list

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced AR-VLA, a standalone autoregressive Action Expert for vision-language-action robots.

  2. It keeps its own long-term history, refreshes vision-language inputs, and uses re-anchoring to reduce the impact of stale perception.

  3. In simulated and real robot tests, it produced smoother actions and matched or improved success rates over reactive VLA baselines.

  4. The approach separates fast control from slower perception and reasoning, offering a more scalable path for robot policy training and deployment.

Read the original