Qwen-Audio-3.0-TTS: From Speaking to Expressive Voice Synthesis
Qwen-Audio-3.0-TTS launches two variants—Flash with ~300 ms latency and Plus with higher naturalness—offering multilingual support for 16 languages, superior WER/CER and speaker similarity scores, free‑style natural‑language control, fine‑grained tag editing, and robust performance in noisy environments, all backed by benchmark results that crown Plus as the top performer on the Artificial Analysis leaderboard.
Qwen-Audio-3.0-TTS is a real‑time speech synthesis model released in two variants: Flash (≈300 ms first‑packet latency) and Plus (higher naturalness and timbre fidelity).
The update focuses on four developer‑pain points: broader language coverage, more natural instruction control, finer‑grained label control, and robustness when reference audio is unclear.
Benchmark on the CV3‑Eval multilingual suite shows the model covers 16 languages (including Chinese, English, Japanese, Korean, German) with 10 languages achieving the best WER/CER. Flash attains an average WER/CER of 3.87, Plus 3.96, both outperforming competing systems.
Speaker similarity scores (higher is better) rank Plus first in all 16 languages with an average SS of 82.75, while Flash follows at 80.44, indicating strong and robust timbre reproduction.
Core highlight 1: comprehensive multilingual and dialect support – 20 Chinese dialects and 16 languages, mitigating the “dialect feature weakening” issue through higher‑quality dialect data and dedicated training.
Core highlight 2: free‑style natural‑language instruction control, allowing users to specify emotion, role, scene, speed, etc., without acoustic expertise.
Core highlight 3: fine‑grained structured tags (e.g., [gasp], [giggles], [angry]) enable precise manipulation of breath and emotion.
Core highlight 4: acoustic robustness in complex environments – specialized training injects speech‑enhancement capability, improving clarity in high‑reverb and high‑noise scenarios.
Additional features include a premium voice library with 20 dialect voices and 14+ minority language voices, 48 kHz high‑definition output (up from 24 kHz), and up to 3‑minute single‑utterance synthesis.
The model is now publicly available; developers are invited to integrate and provide feedback via the technical community.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Alibaba Cloud Developer
Alibaba's official tech channel, featuring all of its technology innovations.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
