Tencent Hunyuan Hy ASR 3.0 Preview Raises Accuracy, Dialect Support, and Robustness
On August 4, Tencent released the Hy ASR 3.0 preview, a speech‑recognition model built on the Hy3 large‑language model that combines a MoE architecture, massive unsupervised audio training and multi‑stage reinforcement learning to cut word‑error rates to around 3 % across Mandarin, English and Cantonese, while improving dialect coverage, context understanding and stability in noisy environments, and is now offered as a cloud API.
On August 4, Tencent officially launched the Hy ASR 3.0 preview, a next‑generation speech‑recognition model that leverages the language‑understanding power of the Hy3 large‑language model.
Evaluations on multiple open‑source test sets show that the model keeps the word‑error rate (WER) near 3 % for all languages, with concrete results of 3.34 % for Mandarin, 2.62 % for English and 3.12 % for Cantonese, outperforming competing products.
On a self‑built evaluation suite that tests general transcription, dialect recognition, context understanding, specialized terminology, and challenging acoustic conditions (high noise, whispering), Hy ASR 3.0 preview also maintains a low WER and leads the competition.
From a user perspective, the model improves four key capabilities:
More accurate universal recognition : higher accuracy for standard speech, dialects, and mixed Mandarin‑English utterances, reducing misspellings, omissions and error accumulation in long audio.
Better understanding of user intent : context‑aware transcription that corrects homophones and eliminates semantic ambiguity.
Easier adaptation to professional scenarios : hot‑word injection to quickly recognize brand names, personal names and industry terms, lowering integration and maintenance costs.
Greater stability in complex environments : dedicated optimizations for high‑noise, whisper, and low‑volume conditions keep performance steady.
The performance gains stem from a combination of architecture, data scaling, and post‑training enhancements. The model adopts an efficiency‑oriented Mixture‑of‑Experts (MoE) architecture and upgrades its base to Hy3, boosting language understanding, context modeling and semantic reasoning. On the acoustic side, a self‑developed unsupervised speech encoder is trained on tens of millions of hours of audio, extracting high‑quality acoustic representations.
To further raise speech‑modeling ability, the team jointly trains the speech encoder with the large‑language model using multi‑source audio data that covers numerous dialects, accents and acoustic scenes, followed by a high‑quality annotation pipeline. After large‑scale joint pre‑training, multi‑stage capability injection gradually endows the model with context awareness, complex‑scene adaptation and dialect recognition.
A high‑quality SFT (Supervised Fine‑Tuning) recipe is built around universal transcription, context understanding, and complex‑scene robustness, covering context, specialized terminology, diverse acoustic environments and varied speaker populations. For dialect recognition, the SFT data set spans ten major dialect regions and more than twenty secondary sub‑regions.
The model also incorporates multi‑stage reinforcement learning targeting general transcription accuracy, any‑context capability, and long‑tail noisy scenarios, reducing mis‑recognition and missed recognitions under difficult acoustic conditions.
Hy ASR 3.0 preview is now publicly available through the Tencent Cloud API portal and can be applied to intelligent customer service, content understanding, voice search and other scenarios. Early adopters such as the Yuanbao product allow users to press‑and‑hold to experience dialect recognition, context‑aware correction and stable transcription in noisy environments, with free access; other products like WorkBuddy are being integrated as well.
API documentation: https://cloud.tencent.com/document/product/1093/135476
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
