Tagged articles

speaker similarity

2 articles · Page 1 of 1
Alibaba Cloud Developer
Alibaba Cloud Developer
Jul 21, 2026 · Artificial Intelligence

Qwen-Audio-3.0-TTS: From Speaking to Expressive Voice Synthesis

Qwen-Audio-3.0-TTS launches two variants—Flash with ~300 ms latency and Plus with higher naturalness—offering multilingual support for 16 languages, superior WER/CER and speaker similarity scores, free‑style natural‑language control, fine‑grained tag editing, and robust performance in noisy environments, all backed by benchmark results that crown Plus as the top performer on the Artificial Analysis leaderboard.

AI modelQwen-Audio-3.0-TTSReal-time inference
0 likes · 7 min read
Qwen-Audio-3.0-TTS: From Speaking to Expressive Voice Synthesis
Weekly Large Model Application
Weekly Large Model Application
Jun 16, 2026 · Artificial Intelligence

Building an Open‑Source TTS Evaluation Framework with ZipVoice, OmniVoice & Latest Benchmarks

This guide explains why TTS evaluation requires a three‑metric “iron triangle” (WER/CER, speaker similarity, and naturalness), introduces community benchmarks such as Seed‑TTS‑eval, TTSDS2, TTS Arena and TTSD‑eval, and provides a concrete six‑stage pipeline and best‑practice checklist for reproducible, production‑ready assessment.

CI pipelineOpen-source benchmarksSeed-TTS-eval
0 likes · 11 min read
Building an Open‑Source TTS Evaluation Framework with ZipVoice, OmniVoice & Latest Benchmarks