Hands-On Test: Qwen-Audio-3.1-TTS-Next Powers Open-Source One-Person Audio Studio
The author tests Alibaba's new Qwen-Audio-3.1-TTS-Next model across diverse scenarios — suspense drama, game NPCs, ads, podcasts — demonstrating its end-to-end generation of layered soundscapes, and releases open-source tools Qwen Audio Studio and z-qwen-audio-studio for local audio production.
Conference Announcements and Model Availability
At the Alibaba Cloud Apsara Conference in Hangzhou, Alibaba unveiled a suite of new Qwen models including Qwen4-Max, Qwen4-27B, Qwen-Image-3.1, Happy Shrimp1.1 (music generation), HappyOyster 2.0 Preview (world model), Qwen-Audio-3.1, Qwen-Audio-3.1-ASR-Next, Qwen-Audio-3.1-TTS-Next , Qwen-Audio-3.1-Realtime, and Qwen3.8-LiveTranslate. The author discovered Qwen-Audio-3.1-TTS-Next already available on the Bailian platform.
Qwen Audio Studio: Open-Source Local Audio Creation Tool
Finding the online UI and API cumbersome, the author built Qwen Audio Studio, a local desktop application positioned as a "one-person audio studio." It offers seven creation modes — podcast, advertisement, audiobook, radio drama, game dubbing, narration, and a custom mode — each automatically loading tailored prompts, instructions, and templates. The tool supports:
Character dialogue, ambient sounds, action sound effects, background music, and reference timbre
Adjustable output format, sample rate, channels, speech rate, volume, and random seed
Instant preview, download, and version comparison
The project is open-sourced on GitHub.
Agent Skill: z-qwen-audio-studio
An automated Agent Skill named z-qwen-audio-studio was also released. Given a natural-language description of a desired podcast, ad, radio drama, or game dub, it structures the prompt (characters, dialogue, ambience, effects, music) and invokes qwen-audio-3.1-tts-next to produce the final audio. This skill is also open-sourced.
Qwen-Audio-3.1-TTS-Next Capability Overview
The model accepts a script or action-annotated text and directly outputs a complete production blending character dialogue, action sound effects, ambient noise, and background music, with precise control over the sequencing of each sound element.
Scenario 1: Suspense Short Drama
Test prompt (convenience store on a rainy night):
测试提示词(Prompt):
广播短剧,电影质感,场景为暴雨深夜的街角便利店
环境声:外头雷雨交加,雨水重重砸在玻璃窗上
门被推开,伴随清脆的进店电子叮咚声,外头的雨声瞬间涌入,随后是收伞甩水的哗啦水滴声,门迅速关上,雨声明显被隔绝在门外,变成闷雷低鸣
店员(二十岁出头男性,嗓音透着熬夜的沙哑与疲惫,有气无力):欢迎光临,关东煮卖完了啊
顾客(五十岁左右中年男性,浑厚、低沉、极其平静,语速极慢):给我一份昨天的报纸
便利店背景里只有冰柜压缩机的微弱蜂鸣,店员敲击键盘的动作突然停顿两秒,吸了一口气道:大叔,您已经连续来买三十年了
紧接着窗外劈过一道沉闷的远雷This scenario is deliberately challenging: it demands accurate acoustic physics — the sharp decay of rain sound as the door opens and closes, the dynamic wet-umbrella shake, the electronic chime, the clerk's weary greeting, and the critical suspense line emerging exactly when the ambient noise yields space.
Scenario 2: Indie Game NPC Interaction
Game developers traditionally rely on dry voice-only assets or middleware-triggered layers in-engine. The test asks the model to fuse an NPC's lines with a living open-world soundscape:
测试提示词(Prompt):
游戏 NPC 台词配音,场景为微风和煦的村口大树下,古槐树叶沙沙作响
背景声:远处隐约的乡村公鸡打鸣与犬吠声、微弱溪流潺潺、远处孩童嬉闹声
NPC 老村长(70 岁男性,嗓音沧桑干哑、语速迟缓慈祥,透着深深的忧虑,伴随拐杖在泥地上轻微的顿地声):孩子啊,你来得正好。后山那边……昨夜又闹起了怪动静,大伙儿心里都慌得很。你本事大,替老汉我走一趟瞧瞧,村里老少,都指望你了。Similar multi-layer scenes were tested extensively, including "Rainy Night Last Bus" (three-character dialogue blended with rain), "Rain Pavilion Return Umbrella" (classical two-person dialogue with rain and wind chimes), and "Snowfield Mechanic" (NPC quest delivery with wind and mechanical sounds).
Production-Driven Tests: Coffee Ad and Podcast
20-Second Premium Coffee Shop Ad
Commercial audio requires gear-like meshing of sound elements: grinder motor, high-pressure steam hiss, coffee drip, porcelain cup on wood, and a poised female narrator entering at the perfect moment.
测试提示词(Prompt):
20 秒精品咖啡品牌声音广告,开场响起咖啡磨豆机的电动运转声,随后是咖啡机高压萃取蒸汽的微弱呲呲声,紧接着一声清脆温润的瓷杯放在木桌上的轻响
广告旁白(三十岁左右知性女性,温暖从容,语速舒缓):世界很吵,但这一刻,是属于你的安静
背景随之缓缓浮现轻柔舒缓的爵士钢琴弱音琶音,旁白轻声念出品牌:木雀咖啡,慢下来,闻见生活Client Revision Workflow
When the client requested a shift from modern street cafe to an old-bookstore cafe — keeping script and narrator timbre but swapping background to creaking wooden floors and soft page turns — the author simply modified the prompt and regenerated, demonstrating rapid iteration.
Late-Night Healing Podcast
Podcasts live or die by "breathing sense" and "companionship sense." The test targets a 28-year-old male host with a low, gentle, slowed delivery, subtle sighs, and natural pauses, layered over a soft late-night ambient bed with faint rain and warm lo-fi music.
测试提示词(Prompt):
情感播客节目,深夜陪伴风格。主播(28 岁男性,嗓音低沉温柔、语速放缓、带有一种深夜炉火般的安抚感)
背景声:舒缓柔和的深夜环境白噪音,带有窗外极微弱的细雨淅沥声与柔和的电台暖色轻音乐
主播台词:“现在是深夜十一点整。如果你这会儿还没睡,希望这段声音能陪你坐一会儿。今天无论遇到了什么糟心事,到这儿都翻篇了。把肩膀松下来,晚安,做个好梦。”(台词中间带有极微弱的叹息声与自然的停顿呼吸)Qwen-Audio-3.1 Family: Real-Time Full-Duplex Interaction
The Qwen-Audio-3.1 family introduces Qwen-Audio-3.1-Realtime , addressing the classic pain points of voice interaction: premature interruption during user pauses and false triggers from background noise. The new model listens better, manages turn-taking, understands emotion, can call tools mid-conversation, and supports full-duplex interruption — making the experience closer to a human phone call. These capabilities are already integrated into Qianwen Office, Qoder, QwenNote A2, QwenNote Eva, and Qianwen AI Glasses. The author previously prototyped a smart speaker with Qwen3.5-Omni (referenced in a prior article) and is now converting an M1 Mac mini into a full-duplex personal desktop assistant.
Paradigm Shift: From Pipeline to End-to-End Generation
Traditional immersive audio production follows a heavy pipeline: script → voice actors (dry recording) → sound-effect library search → multi-track DAW editing → alignment of breaths, EQ, reverb, side-chain ducking — requiring host, sound designer, mixer, and editor. Earlier AI tools merely stitched ASR, single-speaker TTS, and stock libraries, yielding fragile interfaces and unrealistic sound fields.
Qwen-Audio-3.1-TTS-Next adopts a native unified end-to-end generation framework that fuses voice, effects, and ambience into a single physical sound field. A script in; a world with breathing, space, and layers out. This flattens the creation barrier, freeing creators from multi-track manual labor to focus on narrative and creative direction — effectively giving everyone a "one-person audio studio."
Open-Source Resources
Qwen Audio Studio— local audio creation application z-qwen-audio-studio — Agent Skill for automated production
Both repositories (code and prompt templates) are available on GitHub.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Old Zhang's AI Learning
AI practitioner specializing in large-model evaluation and on-premise deployment, agents, AI programming, Vibe Coding, general AI, and broader tech trends, with daily original technical articles.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
