Alibaba Qwen Team Releases Qwen3.5 Omni: A Native Multimodal Model for Text, Audio, Video, and Realtime Interaction

TL;DR AI
2 min readKey summary
Qwen3.5-Omni uses a Thinker-Talker architecture.
It includes a native Audio Transformer (AuT) encoder pre-trained on more than 100 million hours of audio-visual data.
Both the Thinker and the Talker leverage Hybrid-Attention Mixture-of-Experts (MoE).
The architecture supports a 256k long-context input and can ingest over 10 hours of continuous audio.
The article also mentions "over 400 seconds" related to audio/video, but that line is incomplete in the provided content.



