← AI News

Qwen-Audio-3.0-TTS: 16 Languages, Fine-Grained Control over Whisper/Anger/Laughs, #1 on TTS Leaderboard

LaunchSource: xAuthor: Alibaba_QwenHotness: 1238Published Jul 23, 2026

Alibaba released Qwen-Audio-3.0-TTS, the latest text-to-speech model with Flash (real-time) and Plus (high-quality) variants. It features fine-grained inline tags for whisper, anger, breaths, and laughs, natural-language control (e.g., 'read this slowly like a bedtime story'), 16 languages, and up to 3-minute one-pass generation. It is now #1 on the Artificial Analysis TTS Leaderboard.

  • Qwen
  • TTS
  • audio
  • leaderboard
  • Speech Synthesis
View source →

Comments

Log in to comment

No comments yet. Be the first.