Qwen-Audio-3.0-TTS: From Speaking to Expressive Voice Synthesis

Qwen-Audio-3.0-TTS launches two variants—Flash with ~300 ms latency and Plus with higher naturalness—offering multilingual support for 16 languages, superior WER/CER and speaker similarity scores, free‑style natural‑language control, fine‑grained tag editing, and robust performance in noisy environments, all backed by benchmark results that crown Plus as the top performer on the Artificial Analysis leaderboard.

Alibaba Cloud Developer
Alibaba Cloud Developer
Alibaba Cloud Developer
Qwen-Audio-3.0-TTS: From Speaking to Expressive Voice Synthesis

Qwen-Audio-3.0-TTS is a real‑time speech synthesis model released in two variants: Flash (≈300 ms first‑packet latency) and Plus (higher naturalness and timbre fidelity).

The update focuses on four developer‑pain points: broader language coverage, more natural instruction control, finer‑grained label control, and robustness when reference audio is unclear.

Benchmark on the CV3‑Eval multilingual suite shows the model covers 16 languages (including Chinese, English, Japanese, Korean, German) with 10 languages achieving the best WER/CER. Flash attains an average WER/CER of 3.87, Plus 3.96, both outperforming competing systems.

Speaker similarity scores (higher is better) rank Plus first in all 16 languages with an average SS of 82.75, while Flash follows at 80.44, indicating strong and robust timbre reproduction.

Core highlight 1: comprehensive multilingual and dialect support – 20 Chinese dialects and 16 languages, mitigating the “dialect feature weakening” issue through higher‑quality dialect data and dedicated training.

Core highlight 2: free‑style natural‑language instruction control, allowing users to specify emotion, role, scene, speed, etc., without acoustic expertise.

Core highlight 3: fine‑grained structured tags (e.g., [gasp], [giggles], [angry]) enable precise manipulation of breath and emotion.

Core highlight 4: acoustic robustness in complex environments – specialized training injects speech‑enhancement capability, improving clarity in high‑reverb and high‑noise scenarios.

Additional features include a premium voice library with 20 dialect voices and 14+ minority language voices, 48 kHz high‑definition output (up from 24 kHz), and up to 3‑minute single‑utterance synthesis.

The model is now publicly available; developers are invited to integrate and provide feedback via the technical community.

Model performance chart
Model performance chart
Speaker similarity chart
Speaker similarity chart
Fine‑grained tag example
Fine‑grained tag example
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI modelspeech synthesistext-to-speechreal-time inferencemultilingual TTSspeaker similarityQwen-Audio-3.0-TTS
Alibaba Cloud Developer
Written by

Alibaba Cloud Developer

Alibaba's official tech channel, featuring all of its technology innovations.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.