ViDS: Video Diffusion Shader using 3D Face Tracking

TL;DR AI
2 min readKey summary
ViDS is a new portrait animation system that reconstructs a subject-specific 3D face model from a reference image and drives it with expression and pose from video.
It feeds dense geometric cues into a diffusion model to produce more consistent, lifelike talking-head animations.
The paper also adds autoregressive sampling to extend generation beyond the usual window while reducing clip-to-clip artifacts.
Overall, the approach improves realistic portrait animation by using 3D face tracking for better motion control and stronger identity preservation.
