Xiaomi’s Voice AI Chief Daniel Povey Named ISCA Fellow for Pioneering Open‑Source Speech Tools

Daniel Povey, Xiaomi’s chief scientist for speech, was elected a 2026 ISCA Fellow in recognition of his groundbreaking work on acoustic modeling, sequence‑training methods such as Minimum Phone Error, and the creation of the Kaldi toolkit and its next‑generation extensions that power modern multilingual ASR and TTS systems.

Xiaomi Tech
Xiaomi Tech
Xiaomi Tech
Xiaomi’s Voice AI Chief Daniel Povey Named ISCA Fellow for Pioneering Open‑Source Speech Tools

The International Speech Communication Association (ISCA) announced its 2026 Fellows, naming Daniel Povey—Xiaomi Group’s chief scientist for speech—as one of only eight global inductees. ISCA Fellow is the organization’s highest honor, awarded to scholars who have made lasting, distinguished contributions to speech science, technology, and engineering.

Povey earned his Ph.D. at Cambridge University in 2003, then spent roughly a decade at IBM’s T.J. Watson Research Center and Microsoft Research, focusing on core speech‑recognition technologies. He later joined Johns Hopkins University, where he created the Kaldi open‑source speech‑recognition toolkit and co‑authored the LibriSpeech corpus, both of which have become de‑facto standards in academic and industrial research.

Early in his career, Povey introduced the Minimum Phone Error (MPE) sequence‑training method, a pioneering approach that brought discriminative training to speech recognition and laid the groundwork for later techniques such as LF‑MMI. He also championed deep‑learning‑based speaker verification, proposing the X‑vectors method that remains widely used today; his Google Scholar citations exceed 60,000.

In November 2019 Povey moved to Beijing to lead Xiaomi’s AI Lab as chief scientist for speech. Under his leadership the team released the next‑generation Kaldi ecosystem—comprising k2, Lhotse, Icefall, and Sherpa—delivering several notable advances:

Zipformer & Zapformer encoders : Zipformer surpassed Google’s Conformer and was accepted as an oral paper at ICLR 2024 (top 1.2% of submissions); the 2026 Zapformer upgrade improves recognition accuracy by 10‑15 % and enhances training stability and generalization.

OmniVoice multilingual TTS : A zero‑shot synthesis model (0.8 B parameters, initialized from Qwen‑3) supports 600+ languages/dialects, clones voice from 3‑10 seconds of audio, and achieves state‑of‑the‑art results on Chinese/English Seed‑TTS and LibriSpeech‑PC benchmarks, surpassing commercial models such as MiniMax and ElevenLabs.

CR‑CTC (Consistency‑Regularized CTC) : Accepted at ICLR 2025, this method brings pure‑CTC performance on par with Transducer models across LibriSpeech, AISHELL‑1, and GigaSpeech, lowering the barrier for ASR training and deployment.

ZipVoice zero‑shot TTS : Built on Flow Matching and the Zipformer backbone, ZipVoice (single‑speaker) and ZipVoice‑Dialog (dialogue) deliver lightweight, fast inference; a CPU‑single‑thread can run in real time.

sherpa‑onnx inference engine : Supports major NPUs (Qualcomm, Ascend, Rockchip) and platforms (Android, iOS, embedded, server), earning the nickname “Swiss‑army knife of intelligent speech.”

The team’s work has been presented at top conferences such as Interspeech, ICASSP, ICLR, and ACL, resulting in dozens of papers and establishing the new Kaldi suite as core infrastructure for global AI‑speech development.

All components are released under the Apache‑2.0 license, with no commercial usage restrictions, embodying a philosophy of open‑source openness that lowers entry barriers for researchers and developers worldwide. Xiaomi positions the project as public infrastructure rather than an internal tool, enabling speech‑enabled experiences across smartphones, speakers, TVs, cars, and IoT devices, and fostering applications in smart homes, agriculture, accessibility, and international e‑commerce.

Overall, Povey’s election as an ISCA Fellow highlights the impact of open‑source research on both academia and industry, and underscores the role of the next‑generation Kaldi ecosystem in advancing multilingual, low‑resource, and real‑time speech technologies.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

open sourceTTSspeech recognitionASRKaldiDaniel PoveyISCA Fellow
Xiaomi Tech
Written by

Xiaomi Tech

Chat about technology with Xiaomi and change life together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.