Switch language한국어
Back to the list

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model

TL;DR AI

Key summary

2 min read
  1. Researchers introduced CoTinyVLA, a compact 0.9B vision-language-action model built on Qwen3.5-0.8B.

  2. It uses dual-view temporal inputs, hierarchical chain-of-thought distillation from a 35B teacher, and paraphrase augmentation.

  3. The model improves robustness and outperforms 3B- to 7B-parameter baselines on LIBERO-Plus and related benchmarks.

  4. Despite its strong results, CoTinyVLA keeps GPU memory use low, making it attractive for embedded robotics and closed-loop control.

Read the original