ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

TL;DR AI
2 min readKey summary
Researchers introduced ViTacWorld, an action-conditioned world model that fuses vision and touch to predict rollout trajectories for contact-rich robot manipulation.
The model is pretrained on large real and simulated datasets, then fine-tuned with real robot rollouts to better handle physical contact.
ViTacWorld can support data augmentation and policy evaluation, helping reduce the need for costly real-world tactile data collection.
The work points to a scalable path for improving simulation-to-real manipulation with more reliable visuo-tactile learning.
