Switch language한국어
Back to the list

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

TL;DR AI

Key summary

2 min read
  1. CoInteract is an end-to-end diffusion-based video generation framework conditioned on reference images, text, and audio.

  2. It uses a human-aware mixture-of-experts and a dual-stream co-generation scheme to better preserve local detail and interaction geometry.

  3. The approach improves structural stability and contact realism in human-object interaction videos.

  4. That makes it especially relevant for advertising, e-commerce, and virtual marketing workflows that need physically believable video.

Read the original