Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

TL;DR AI
2 min readKey summary
Alibaba’s Tongyi Lab launched Qwen-Audio-3.0-TTS as a hosted text-to-speech service in Flash and Plus tiers.
The model supports 16 languages, streaming APIs over WebSocket, voice cloning, and detailed style controls via inline tags.
Flash targets low latency, while Plus is tuned for higher quality and currently ranks first on an independent TTS leaderboard.
The release gives developers production-ready TTS without self-hosting model weights, while improving robustness and speech control.

