TaoMate-H3 Open-Sourced: 3-Step LoRA Streaming for Minute-Long Audio-Video Generation

Alibaba's TaoLive AIGC team open-sources TaoMate-H3, a streaming audio-video generation model based on MiniMax H3 that uses three-step LoRA inference and autoregressive generation to produce minute-long videos with synchronized dialogue, singing, and sound effects at multiple resolutions, achieving 11-27x speedups on H20 GPUs.

Machine Heart
Machine Heart
Machine Heart
TaoMate-H3 Open-Sourced: 3-Step LoRA Streaming for Minute-Long Audio-Video Generation

Overview

TaoMate-H3 is an open-source audio-video joint streaming generation model developed by Alibaba's TaoLive AIGC team, based on MiniMax H3. It combines three-step LoRA inference with autoregressive generation to continuously generate videos containing dialogue, singing, or environmental sounds from text descriptions, supporting minute-level continuation and 480p, 768p, 1080p horizontal/vertical output.

Core Capabilities

Three-step streaming generation: Each small segment requires only three denoising steps, significantly reducing first-segment latency and enabling generate-while-playing for live streaming, virtual characters, and interactive video.

Audio-video joint generation: Handles both sound and visuals in a single process, supporting character speech, singing, dialogue, and environmental sounds (waves, vehicles, explosions). Audio participates in generation to synchronize lip movements, expressions, and actions.

Minute-level continuous continuation: Uses continuously updated audio-video KV cache and segment-level audio guidance to extend content across prompt boundaries, allowing creators to adjust focus per segment for complete product introductions or character narratives.

Application Demonstrations

The release includes several demo categories:

E-commerce explanation: One-minute vertical backpack introduction and horizontal wooden comb explanation, showing different aspect ratios and complete generation results. Creators can tailor explanations for different audiences (commuting, travel, daily use) with segment-level control.

Singing and character performance: "Window singing" demo shows a realistic character singing with melody, lip-sync, and body motion coordinated.

Anime dialogue: "Moonlit girl and white fox" demonstrates stylized characters, environment, and dialogue in a short narrative segment.

Environment and effects: "Sunset coast with sea erosion arch" uses waves and ambient sound for atmosphere; "Character walking forward, building exploding behind" combines character motion, building explosion, and impact sound for intense audiovisual rhythm.

All demos cover multiple content types and support both aspect ratios and multiple resolutions for mobile and widescreen adaptation.

Overall Inference Pipeline

From prompt to final video, TaoMate-H3's inference consists of four stages:

Segment content organization: Reads prompt sequence according to target duration, determining scene, action, and sound description for each segment.

Audio guidance preparation: Generates audio denoising trajectory for current segment, using previous segment's tail audio to connect timbre and tone.

Joint streaming generation: Splits segment into small chunks, each executing three-step audio-video denoising while continuously updating historical KV cache.

Decoding output: Generated latents are decoded via VAE into frames and audio; the public inference entry runs the full pipeline and saves the final video.

Latent refers to the compressed audio-video representation used during generation; the model operates in this space before decoding to watchable/listenable content.

Key Technologies

Three-Step LoRA: Reducing Per-Chunk Generation Overhead

TaoMate-H3 adapts few-step LoRA to compress streaming inference into three denoising intervals. Under the current 5-second prompt configuration, each segment is divided into four chunks: main chunk ~1.4 seconds, tail chunks shorter. The model completes chunks sequentially, using historical context to continue generation. The three steps advance along time states 0→16→33→49. After each chunk, the model updates the KV cache based on the final generated audio-video state for the next chunk. Few-step generation reduces repeated iteration compute; chunked processing allows earlier first-segment output.

Self Forcing: Autoregressive Training for Continuous Continuation

During training, TaoMate-H3 employs Self Forcing, letting the model condition on its own generated history chunks to produce subsequent content. This exposes training to actual continuation-time history states, mitigating the train-inference distribution gap from relying solely on ground-truth history. This training approach, combined with few-step inference and KV cache, forms the foundation of streaming continuation: the model learns both current-chunk generation and how to leverage prior outputs to continue characters, actions, and scenes.

Reference: https://arxiv.org/pdf/2506.08009
Self Forcing training diagram
Self Forcing training diagram

Audio-Video KV Cache: Supporting Long-Video Continuous Memory

To maintain long-video consistency, TaoMate-H3 continuously reuses audio-video KV cache, passing generated context to subsequent chunks. Character state, recent actions, and audio information thus persist over time, remaining active even when prompts switch. The cache uses a "initial visual reference + recent audio-video context" organization: the initial reference preserves character and scene initialization, while recent context rolls forward with generation progress, letting the model follow the latest actions and sounds. This bounds history cache at a fixed size, avoiding unbounded growth with video length, enabling stable minute-level continuous generation.

Continuous Time Encoding: Improving Cross-Prompt Transitions

In long videos, prompts may change while the media timeline must stay continuous. TaoMate-H3 uses globally continuous RoPE positional encoding, moving the current text's time coordinate with audio-video progress and aligning the text's right boundary to the current media start, so new prompts correspond to the content being generated. When switching prompts, the system retains existing audio-video KV cache while distinguishing newly generated content from the tail context needed for encoding/decoding. Historical content is only used for bridging, not re-generated. To address brightness and color drift in long continuations, the model references the first chunk's visual statistics to align mean and variance of subsequent video latents, reducing perceptual drift.

Audio Trajectory Guidance: Stabilizing Dialogue Rhythm and Timbre

Streaming generation splits video into short chunks, but a full line of dialogue requires pauses, speech rate, and ending position arranged within the complete segment. TaoMate-H3 first prepares an audio denoising trajectory covering the current prompt segment, then uses the corresponding time-position audio latents to guide each small chunk, ensuring local generation always has segment-level audio reference. The three denoising updates align with three states in the audio trajectory, and the final audio latent matches the clean endpoint of the guidance trajectory. Audio guidance runs throughout generation, stabilizing dialogue pacing and reducing repeated sentence starts in short-chunk continuation. To bridge timbre and tone across segments, the system uses the last ~1 second of clean audio from the previous segment as a read-only reference; new audio continues under this reference, which remains unchanged; finally all segment audio latents are decoded together for continuous sound output.

Performance Evaluation

On a single 8×NVIDIA H20 96 GB machine, the team compared vanilla MiniMax H3 with TaoMate-H3, both outputting 480×864, 10-second video.

Pure DiT compute time: reduced from 169.572 seconds to 14.810 seconds ( 11.45× speedup ).

First-segment final latent ready time: reduced from 170.052 seconds to 6.148 seconds ( 27.66× speedup ). This corresponds to the time when the model finishes the first segment and can hand off to VAE decoding.

Performance comparison chart
Performance comparison chart

Under this 10-second configuration, TaoMate-H3 generates 8 small chunks, executing 24 forward passes and 8 KV updates. Single-node 8-GPU uses TP2×Ulysses4 parallelism, combined with efficient attention, operator fusion, and selective low-precision computation for some linear layers, reducing inference overhead. These optimizations improve both full-sequence generation efficiency and first-segment latency, providing a foundation for chunked decoding, playback, and interactive application development.

Quick Start

Environment and Hardware

Supports Linux, Python 3.10/3.11, CUDA 12.8, PyTorch 2.8, with single-node 4-GPU or 8-GPU inference entry points. Verified on 8×H20 96 GB; full installation steps in repository README.

Run Example

git clone https://github.com/TaoLiveAIGC/TaoMate-H3.git
cd TaoMate-H3

After completing environment setup and downloading MiniMax H3 base weights per README, run the 10-second example. TaoMate-H3 LoRA weights download automatically on first use:

python -m taomate_h3 \
  --model-root models/MiniMax-H3 \
  --prompt-json examples/prompts_10s.json \
  --duration 10 \
  --resolution 768x1376 \
  --gpus 8 \
  --output outputs/demo_10s

Output saved to outputs/demo_10s/video.mp4. Replace the prompt file to change scene, dialogue, and actions; adjust duration and resolution parameters to generate longer videos or switch aspect ratios.

Open Source and Future Plans

This release includes inference code, LoRA weights, and example prompts. Developers are welcome to download and experiment, and to engage on streaming inference, audio-video generation, and application integration. The project is based on MiniMax H3 and follows the MiniMax H3 Community License. Thanks to MiniMax H3 and related open-source contributors.

The team continues to optimize inference speed, further reducing first-segment and subsequent response wait times. They also plan to soon release and open-source a new FL2AV streaming model, further expanding to Ref2AV, with corresponding weights opened, supporting richer audio-video creation workflows. The goal is to advance audio-video streaming generation with the community and bring this technology into more practical creative scenarios.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LoRAGPU inferenceopen sourceautoregressiveaudio-video generationstreaming generationMiniMax H3TaoMate-H3
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.