Tagged articles

multimodal inference

3 articles · Page 1 of 1
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Aug 5, 2026 · Artificial Intelligence

Accelerating Xiaohongshu ‘Ask’ Inference: Slim Vision Tokens, Focused MoE

The article details how Xiaohongshu’s multimodal “Ask” service reduces visual token bloat and MoE expert compute by applying dynamic Vision Token compression (VisionZip) and similarity‑based expert re‑routing (SERE), achieving up to 13% lower first‑token latency and 16.5% faster end‑to‑end throughput while preserving answer quality.

MoE expert routingQwen modelsSERE
0 likes · 19 min read
Accelerating Xiaohongshu ‘Ask’ Inference: Slim Vision Tokens, Focused MoE
Machine Heart
Machine Heart
Jun 2, 2026 · Artificial Intelligence

WorldCache Boosts Video World Model Inference Up to 3.7× with Near‑Lossless Quality

WorldCache separates cacheable and recomputable tokens in diffusion world models using curvature‑based classification and a chaotic‑prioritized adaptive skipping schedule, achieving up to 3.7× speedup on HunyuanVoyager‑13B and Aether‑5B without extra memory or retraining while preserving visual quality.

AI accelerationWorldCacheadaptive skipping
0 likes · 7 min read
WorldCache Boosts Video World Model Inference Up to 3.7× with Near‑Lossless Quality