Data Party THU
Aug 2, 2026 · Artificial Intelligence
Masked Visual Actions Enable Generalizable Robot Modeling via Pixel Trajectories
The paper introduces Masked Visual Actions, a pixel‑mask representation of robot behavior that lets a 14B video model predict future outcomes and generate robot motions across unseen embodiments, achieving higher accuracy than traditional joint‑angle or pose inputs.
Roboticscross-embodiment generalizationmasked visual actions
0 likes · 9 min read
