Xiaomi‑Robotics‑U0: An Open‑Source 38B Multimodal Model that Scales Embodied Data Generation

Xiaomi‑Robotics‑U0, a 38‑billion‑parameter multimodal autoregressive model, unifies four embodied generation tasks, offers geometry‑preserving data augmentation, achieves an 83× inference speedup with FlashAR+, tops the WorldArena benchmark, and improves real‑robot OOD task completion by over 26%.

Xiaomi Tech
Xiaomi Tech
Xiaomi Tech
Xiaomi‑Robotics‑U0: An Open‑Source 38B Multimodal Model that Scales Embodied Data Generation

Today Xiaomi officially released Xiaomi‑Robotics‑U0 , a 38 billion‑parameter multimodal autoregressive foundation model that unifies four core embodied generation tasks: scene generation, embodied transfer, video generation, and text‑to‑image/anything‑to‑image editing.

The model can augment existing robot data—changing objects, lighting, background, or adding interference—while preserving geometric consistency, eliminating the need for costly re‑collection. Using the FlashAR+ inference acceleration scheme, generation efficiency is increased by nearly 83 times compared with the original autoregressive paradigm.

On the WorldArena benchmark, which evaluated 126 models, Xiaomi‑Robotics‑U0 achieved the overall first‑place score. In real‑robot out‑of‑distribution (OOD) tests, strategies that augment training data with the model improve task‑completion progress by an average of >26 %.

To achieve precise per‑frame control, the authors introduced a five‑dimensional decoupled structured control paradigm. The five dimensions—workbench layout, foreground objects, unrelated foreground clutter, lighting conditions, and background information—can each be independently manipulated via natural language, ensuring multi‑view geometric consistency across different robot platforms.

Compared with the closed‑source GPT‑Image‑2.0 on a 300‑sample difficulty‑balanced test set, Xiaomi‑Robotics‑U0 markedly outperforms in depth consistency, structural fidelity, and semantic alignment, thanks to its fully unified multi‑view geometry.

In video generation, the model ranked first on WorldArena for instruction following, interaction quality, and perspectivity, producing coherent long‑duration interaction videos and supporting virtual camera motion simulation, which is valuable for offline robot strategy validation.

FlashAR+ builds on a multimodal autoregressive architecture that uses IBQ as an image tokenizer, enabling parallel decoding and vLLM‑style KV‑cache paging. Through this, throughput for a 1024×1024 image drops from 450.77 s per image to 5.44 s—a speedup of 82.9 ×—while maintaining generation quality.

All code, model weights, and documentation are fully open‑source. Project homepage, GitHub repository, and HuggingFace model collection are provided for the embodied‑intelligence community.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Data AugmentationMultimodal ModelWorldArena BenchmarkFlashAR+Xiaomi RoboticsEmbodied Generation
Xiaomi Tech
Written by

Xiaomi Tech

Chat about technology with Xiaomi and change life together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.