Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a hosted text-to-speech system available in two variants: Flash for real-time interaction and Plus for high-quality generation. Both tiers are delivered via Alibaba Cloud Model Studio rather than as downloadable weights.

  • The model supports 16 languages and 20 Chinese dialect regions, with Flash achieving the lowest average word/character error rate at 3.87.
  • Qwen-Audio-3.0-TTS-Plus ranks first on the Artificial Analysis Text-to-Speech leaderboard with an Elo score near 1,236.
  • The system includes 86 fine-grained inline tags for non-verbal details like laughter and breathing, alongside natural-language style control.
  • Flash targets a first-packet latency of 300 ms, while Plus focuses on timbre fidelity and speaker similarity, ranking top across all 16 languages.

The release provides developers with robust multilingual synthesis capabilities and fine-grained audio control through a hosted API.