Muka Robotics' LJM Secures WorldArena Runner‑Up Spot with Full‑Process Training on Alibaba Cloud PAI

Muka Robotics' embodied world model LJM achieved second place in the WorldArena leaderboard with an EWMScore_P of 73.06, thanks to a dual‑expert architecture, MoT shared attention, Value‑Driven Temporal Conditioning, and full‑process training on 32 Alibaba Cloud Zhenwu 810E GPUs, which also delivered state‑of‑the‑art results on the LIBERO benchmark.

Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Muka Robotics' LJM Secures WorldArena Runner‑Up Spot with Full‑Process Training on Alibaba Cloud PAI

The Latent Joint‑conditional Model (LJM) developed by Muka Robotics placed second on the public WorldArena leaderboard, scoring 73.06 EWMScore_P—just 0.58 points behind the top‑ranked Xiaomi UNIS (73.64). The model excels in 3D Accuracy (92.60), Motion Quality (89.17) and Controllability (85.93), demonstrating strong spatial cognition and action‑control capabilities.

LJM’s core innovation is a two‑expert design that treats robot actions as interactive latents rather than low‑dimensional numeric inputs. The Latent Reasoning Expert uses Qwen3‑VL‑2B to encode language, visual history, current observations and future actions into a sequence of reasoning tokens that predict future states in latent space. The World Modeling Expert employs the Wan2.2‑TI2V‑5B Diffusion Transformer to generate future video frames in VAE latent space, converting the reasoning constraints into continuous video output.

Both experts exchange information layer‑by‑layer via a Mixture‑of‑Transformers (MoT) shared‑attention mechanism, while Value‑Driven Temporal Conditioning (VTC) preserves high‑information‑density history states, mitigating long‑term memory drift.

On the LIBERO robot‑operation benchmark, LJM outperformed all baselines across PSNR, SSIM, FVD and LPIPS. Compared with the strongest non‑LJM baseline, PSNR improved from 27.25 to 28.60, SSIM from 0.8820 to 0.9238, FVD dropped from 77.09 to 46.49 (≈ 39.7 % reduction), and LPIPS fell from 0.0639 to 0.0247 (≈ 61.3 % reduction), indicating superior frame‑level reconstruction and long‑term video consistency.

Training LJM required handling language, vision and action modalities jointly. The model was pre‑trained on the DROID dataset and fine‑tuned on RoboTwin, processing 65 frames per sample (8 historical, 1 current anchor, 56 future). Alibaba Cloud’s AI platform PAI supplied the compute backbone: 32 Zhenwu 810E GPUs, a global batch size of 32, AdamW optimizer, BF16 mixed precision and ZeRO‑3 parallelism. PAI’s native support for heterogeneous clusters, efficient gradient synchronization, and flexible resource elasticity accelerated the multi‑expert training pipeline and reduced iteration time.

Beyond raw performance, PAI offered end‑to‑end tooling for data preprocessing, multi‑stage training orchestration, monitoring, checkpoint recovery and integration with the NVIDIA Physical AI stack, allowing the Muka Robotics team to focus on algorithmic advances rather than infrastructure management.

The success of LJM illustrates how advanced embodied world models are transitioning from laboratory research to industrial impact, with Chinese teams increasingly leading in both perception and physical reasoning. Continued collaboration between model innovators and cloud providers like Alibaba Cloud is expected to further accelerate this shift.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

embodied AIRoboticsbenchmarkingWorld Modelsmultimodal diffusionAlibaba Cloud PAIlatent joint conditioning
Alibaba Cloud Big Data AI Platform
Written by

Alibaba Cloud Big Data AI Platform

The Alibaba Cloud Big Data AI Platform builds on Alibaba’s leading cloud infrastructure, big‑data and AI engineering capabilities, scenario algorithms, and extensive industry experience to offer enterprises and developers a one‑stop, cloud‑native big‑data and AI capability suite. It boosts AI development efficiency, enables large‑scale AI deployment across industries, and drives business value.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.