Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
TL;DR AI
2 min readKey summary
Researchers introduced SwanSphere, a unified streaming system for high-quality spatial audio generation from panoramic video and text prompts.
It uses a causal autoregressive diffusion transformer, spatial video-audio contrastive learning, online direct preference optimization, and automated spatial captioning.
The system improves both video-to-spatial-audio and text-to-spatial-audio performance in tests.
SwanSphere is designed to reduce latency while improving multimodal alignment and spatial accuracy for immersive media applications.
