Tagged articles

video prediction

2 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 22, 2026 · Artificial Intelligence

Masked Visual Actions: Controlling Robots with Only 15 Hours of Video

A new world model called Masked Visual Actions uses just 15 hours of robot video to predict action outcomes and generate robot behavior by representing motions as spatiotemporal pixel masks, achieving cross‑embodiment generalization, higher task success rates, and strong correlation between video evaluation and real‑world performance.

cross-embodiment generalizationinverse kinematicsmasked visual actions
0 likes · 9 min read
Masked Visual Actions: Controlling Robots with Only 15 Hours of Video
AIWalker
AIWalker
Mar 6, 2025 · Artificial Intelligence

How SCMHSA Improves Transformer Next‑Frame Prediction by Reducing Semantic Dilution

The paper introduces a Semantic‑Concentrated Multi‑Head Self‑Attention (SCMHSA) module and a new embedding‑space loss to address semantic dilution and loss‑target mismatch in Transformer‑based video next‑frame prediction, demonstrating significant PSNR and MSE gains across four benchmark datasets.

Embedding LossSCMHSASemantic Dilution
0 likes · 23 min read
How SCMHSA Improves Transformer Next‑Frame Prediction by Reducing Semantic Dilution