Qwen-Audio-3.0-TTS: 16 Languages, Fine-Grained Control over Whisper/Anger/Laughs, #1 on TTS Leaderboard
LaunchSource: xAuthor: Alibaba_QwenHotness: 1238Published Jul 23, 2026
Alibaba released Qwen-Audio-3.0-TTS, the latest text-to-speech model with Flash (real-time) and Plus (high-quality) variants. It features fine-grained inline tags for whisper, anger, breaths, and laughs, natural-language control (e.g., 'read this slowly like a bedtime story'), 16 languages, and up to 3-minute one-pass generation. It is now #1 on the Artificial Analysis TTS Leaderboard.
- Qwen
- TTS
- audio
- leaderboard
- Speech Synthesis
Comments
Log in to comment
No comments yet. Be the first.