CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
TL;DR AI
2 min readKey summary
CoInteract is an end-to-end diffusion-based video generation framework conditioned on reference images, text, and audio.
It uses a human-aware mixture-of-experts and a dual-stream co-generation scheme to better preserve local detail and interaction geometry.
The approach improves structural stability and contact realism in human-object interaction videos.
That makes it especially relevant for advertising, e-commerce, and virtual marketing workflows that need physically believable video.
