Switch language한국어
Back to the list

Qwen3.5-Omni debuts with text, code, audio, video and web capabilities

TL;DR AI

Key summary

2 min read
  1. Alibaba’s Tongyi Lab announced Qwen3.5-Omni on March 30, 2026 as an omni‑modal model handling text, image, audio and video.

  2. The model was trained on over 100 million hours of audio‑visual data and includes Hybrid MoE Talker and Thinker components for contextual audio generation.

  3. Qwen3.5-Omni supports long inputs (up to 256,000 tokens), 10 hours of audio or 400 seconds of audio‑visual data, and multilingual speech recognition and synthesis.

  4. The release includes Omni Plus, Omni Flash, and Omni Light variants available via offline and real‑time APIs, and benchmarks show Omni Plus outperforming Gemini 3.1 Pro on several tests.

Read the original