Machine Learning Algorithms & Natural Language Processing
Jul 22, 2026 · Artificial Intelligence
Masked Visual Actions: Controlling Robots with Only 15 Hours of Video
A new world model called Masked Visual Actions uses just 15 hours of robot video to predict action outcomes and generate robot behavior by representing motions as spatiotemporal pixel masks, achieving cross‑embodiment generalization, higher task success rates, and strong correlation between video evaluation and real‑world performance.
Roboticscross-embodiment generalizationinverse kinematics
0 likes · 9 min read
