Painting Robot Actions into Video World Models for Bidirectional Inference (Fei‑Fei Li Co‑author)

The paper introduces Masked Visual Actions, a method that renders robot motions as pixel‑level masks for video world models, enabling both forward prediction of environment changes and inverse generation of robot actions, achieving notable success‑rate gains and cross‑embodiment generalization.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Painting Robot Actions into Video World Models for Bidirectional Inference (Fei‑Fei Li Co‑author)

Model‑driven robot planning traditionally requires generating multiple candidate actions, feeding them to a world model for visual preview, and then selecting an executable trajectory. However, learning a mapping between numeric action parameters and visual observations demands extra data and does not transfer well across different robot morphologies.

The recent work Masked Visual Actions for Unified World Modeling (first author Hadi Alzayer, with Fei‑Fei Li, Jia‑Jun Wu, Lu‑Min Zhang, etc.) proposes to render candidate robot motions as pixel‑level motion masks and let a video‑prediction model complete the surrounding environment. This formulation supports two directions: (1) given robot actions, predict environmental response; (2) given desired object motion, generate plausible robot actions.

Training data are constructed in two ways: (1) directly segmenting real robot videos, preserving authentic visual cues; (2) synthetically rendering robot masks from URDF models, camera parameters, and robot states, which allows arbitrary action trajectories at inference time. The model is fine‑tuned from the 14B Wan‑Fun‑Control 2.2 foundation model using LoRA, consuming roughly 15 hours of robot interaction data (≈1 000 DROID demonstrations and 4 000 RoboCasa samples). Training runs for about 10 000 steps on eight H200 GPUs over four days.

Empirically, integrating the video model into planning raises success rates on RoboCasa tasks (e.g., closing a microwave, opening a drawer) by 7–26 percentage points. The predicted trajectories correlate with real‑simulation outcomes at 0.982. In the CoffeeServeMug benchmark, the inverse‑action pipeline—video model generation followed by an inverse‑kinematics extractor—achieves 90 % success, outperforming Diffusion Policy, ACT, and SmolVLA.

Because the full pixel mask encodes robot shape, occupancy, occlusion, and contact boundaries, the approach exhibits strong cross‑embodiment generalization. When tested on unseen grippers or a dual‑arm robot, mask‑based conditioning maintains stable predictions, whereas end‑effector points or skeletal inputs cause the model to hallucinate incorrect robot forms or static scenes.

Limitations include dependence on a large video foundation model, optimistic bias in generated futures, and the need for a separate inverse‑kinematics module to obtain executable commands. Consequently, the method is best suited as a low‑cost, early‑stage planner for strategy development rather than a real‑time control solution.

Figure
Figure
Figure
Figure
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

robotic planningvideo world modelingcross-embodiment generalizationmasked visual actionsinverse action generationsimulation evaluation
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.