QuantWM: Training-Free 2-bit KV Cache Quantization for Stable Video World Models
Researchers from HIT Shenzhen and NUS propose QuantWM, a training-free 2-bit KV Cache quantization framework that reduces temporal flickering in video world models by protecting attention mechanisms, achieving up to 6.20x compression while improving visual quality over existing methods.
Video world models generate continuous frames while maintaining scene consistency under user-specified actions. Autoregressive video world models rely on KV Cache to store historical frames for retrieval and reuse, but the memory cost grows rapidly. For example, LingBot-World-v2 generating ~5 seconds (93 frames) requires over 21 GiB of GPU memory for KV Cache.
Quantizing KV Cache to 2-bit is a promising direction. Existing methods like Quant-VideoGen (QVG) apply 2-bit KV Cache quantization to autoregressive video generation and achieve VBench scores close to BF16 baselines. However, when applied to video world models, these methods still produce noticeable flickering, blurring, and detail loss despite near-identical benchmark scores. Using HY-World 1.5 as an example, BF16 and QVG achieve VBench Temporal Flickering scores of 95.62% and 95.85% respectively, yet visual inspection reveals severe flickering and degradation (Figure 1). This indicates that even flickering-related metrics fail to capture quantization-induced degradation in video world models. The team therefore combined video visualization with per-frame PSNR, SSIM, and LPIPS relative to BF16 outputs to analyze quantization deviations.
The team investigated why Key quantization causes larger visual degradation than Value quantization despite smaller quantization error. Keys act as indices for retrieving historical information, while Values are the retrieved content. Quantizing only Keys to 2-bit caused more pronounced visual degradation than quantizing only Values. Moreover, Keys exhibit smaller quantization error than Values yet induce larger output errors, meaning quantization error magnitude alone cannot explain temporal flickering and quality loss.
Analyzing attention computation in video world models revealed that the matching scores between current Query and historical Keys determine how the model allocates attention to history. Quantization error in Keys can alter the relative ranking of these scores, causing the model to attend to a different frame or spatial location. Figure 2 shows that for the same Query, the highest-scoring historical spatiotemporal position changes after Key quantization, sometimes jumping across frames. For continuous generation, this shift disturbs the model's use of historical scenes, leading to temporal flickering and quality degradation. Therefore, compressing KV Cache must also protect how the model retrieves historical information.
QuantWM Framework
QuantWM introduces two training-free components:
1. Quantization Sensitivity-Aware Clustering (QSAC)
In QVG, each Key is assigned to a cluster center (kept at higher precision) and the residual is quantized to 2-bit. The team observed that the nearest center does not necessarily yield the best post-quantization result because the residual's dynamic range and the channels where error falls affect final attention scores. QSAC first selects candidate centers using a Query-sensitivity-weighted distance, then estimates the attention perturbation that INT2 quantization of each residual group would introduce by considering both dynamic range and sensitivity. The center minimizing estimated perturbation is chosen, thereby reducing impactful errors at the quantization stage.
2. Principal Subspace Attention Compensation (PSAC)
Even after QSAC, 2-bit quantization leaves unavoidable errors. Directly compensating attention logits would require storing full quantization error, which is expensive. The team observed that historical Query energy concentrates in a few principal directions. For instance, in one layer of LingBot-World-v2, the top 8 principal directions cover ~71% of Query energy. PSAC stores the Key quantization error's components in these principal directions in low-rank form and compensates the Key during cache reconstruction, correcting subsequent attention scores. Experiments use 8 principal directions and quantize compensation coefficients to INT8 to control storage overhead. Both QSAC and PSAC rely on statistics from already-generated historical Queries, requiring no model retraining and plugging directly into the video world model generation pipeline.
Experimental Validation
The team evaluated QuantWM on three video world models (Matrix-Game-2, LingBot-World-v2, HY-World 1.5) and two video generation models (LongCat-Video, Causal-Forcing) for generalization. At 480p, 93 frames, QuantWM outperforms QVG and KIVI on PSNR, SSIM, and LPIPS, and visualizations show significantly reduced flickering and blurring (Figure 3). On VBench's five evaluation dimensions, QuantWM achieves the best average rank among compared quantization methods across all models. Additional evaluations at 720p and 1-minute long videos confirm quality at higher resolution and benchmark performance for long-form generation.
Attention analysis confirms QuantWM reduces quantization-induced shifts in historical spatiotemporal position selection. Table 1 shows the proportion of highest-scoring historical tokens that differ from BF16: on Matrix-Game-2 it drops from 53.88% to 10.19%, and on HY-World 1.5 from 61.36% to 13.92%.
System Implementation
QuantWM implements custom Triton kernels. Accounting for cluster centers, quantization parameters, and compensation information, the actual KV Cache compression reaches up to 6.20x (Figure 4).
Conclusion
QuantWM reveals an overlooked issue in low-bit KV Cache compression for video world models: close benchmark scores do not guarantee stable utilization of historical scene information. Therefore, compression research for video world models requires multi-faceted evaluation combining benchmark metrics, actual video quality, and analysis of how the model leverages historical information.
Paper: https://arxiv.org/abs/2609.26425
Project: https://quantwm-project.github.io/QuantWM/
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
