Switch language한국어
Back to the list

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures

TL;DR AI

Key summary

2 min read
  1. A study of three vision-language-action robot policies shows failure signals depend on the model architecture, so monitoring should not be one-size-fits-all.

  2. Across 450 episodes, direction reversal emerged as a strong universal predictor of failure.

  3. Jerk was useful mainly for discrete-token models like VQ-BeT, while velocity checks were often weak or ineffective, especially for continuous policies.

  4. The researchers introduced SafeContract, a training-free black-box monitoring toolkit with conformal calibration to better detect robot policy failures.

Read the original