Xiaomi‑Robotics‑U0: An Open‑Source 38B Multimodal Model that Scales Embodied Data Generation
Xiaomi‑Robotics‑U0, a 38‑billion‑parameter multimodal autoregressive model, unifies four embodied generation tasks, offers geometry‑preserving data augmentation, achieves an 83× inference speedup with FlashAR+, tops the WorldArena benchmark, and improves real‑robot OOD task completion by over 26%.
Today Xiaomi officially released Xiaomi‑Robotics‑U0 , a 38 billion‑parameter multimodal autoregressive foundation model that unifies four core embodied generation tasks: scene generation, embodied transfer, video generation, and text‑to‑image/anything‑to‑image editing.
The model can augment existing robot data—changing objects, lighting, background, or adding interference—while preserving geometric consistency, eliminating the need for costly re‑collection. Using the FlashAR+ inference acceleration scheme, generation efficiency is increased by nearly 83 times compared with the original autoregressive paradigm.
On the WorldArena benchmark, which evaluated 126 models, Xiaomi‑Robotics‑U0 achieved the overall first‑place score. In real‑robot out‑of‑distribution (OOD) tests, strategies that augment training data with the model improve task‑completion progress by an average of >26 %.
To achieve precise per‑frame control, the authors introduced a five‑dimensional decoupled structured control paradigm. The five dimensions—workbench layout, foreground objects, unrelated foreground clutter, lighting conditions, and background information—can each be independently manipulated via natural language, ensuring multi‑view geometric consistency across different robot platforms.
Compared with the closed‑source GPT‑Image‑2.0 on a 300‑sample difficulty‑balanced test set, Xiaomi‑Robotics‑U0 markedly outperforms in depth consistency, structural fidelity, and semantic alignment, thanks to its fully unified multi‑view geometry.
In video generation, the model ranked first on WorldArena for instruction following, interaction quality, and perspectivity, producing coherent long‑duration interaction videos and supporting virtual camera motion simulation, which is valuable for offline robot strategy validation.
FlashAR+ builds on a multimodal autoregressive architecture that uses IBQ as an image tokenizer, enabling parallel decoding and vLLM‑style KV‑cache paging. Through this, throughput for a 1024×1024 image drops from 450.77 s per image to 5.44 s—a speedup of 82.9 ×—while maintaining generation quality.
All code, model weights, and documentation are fully open‑source. Project homepage, GitHub repository, and HuggingFace model collection are provided for the embodied‑intelligence community.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
