OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
TL;DR AI
2 min readKey summary
Researchers introduced OmniVAE, a jointly trained audio-video variational autoencoder for omnimodal generation.
It aligns audio and video latent spaces with segment-level contrastive learning and feature distillation from pretrained semantic encoders.
The goal is to improve synchronized text-to-audio-video generation by learning aligned multimodal representations upfront.
The approach targets a key bottleneck in joint generation: better cross-modal alignment and timing consistency.
