OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation

TL;DR AI
2 min readKey summary
Researchers introduced OASIS, a visuomotor policy for robotic manipulation that fuses vision-language and depth features.
OASIS predicts camera-frame SE(3) end-effector trajectories, then uses pose-aware states to generate action chunks.
The method improves alignment between perception and control, helping robots reason about action-space geometry more directly.
In simulation and real-world tests, OASIS outperformed VLA and world action model baselines, with stronger success and generalization.
