Tagged articles

Video Prediction

3 articles · Page 1 of 1
AntTech
AntTech
Sep 2, 2026 · Artificial Intelligence

ECCV 2026 Paper Showcase: Breaking Vision Perception & Prediction Boundaries

This article summarizes four ECCV 2026 papers: a dynamic cross-layer injection framework for deep vision-language fusion, an event-augmented VLA model enabling robot operation in extreme darkness and blur, a predictive differentiable rendering method using 2D Gaussians for high-fidelity video prediction, and a data-scaling approach for high-resolution weather forecasting that demonstrates clear scaling laws.

Data ScalingECCV 2026Embodied AI
0 likes · 10 min read
ECCV 2026 Paper Showcase: Breaking Vision Perception & Prediction Boundaries
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 22, 2026 · Artificial Intelligence

Masked Visual Actions: Controlling Robots with Only 15 Hours of Video

A new world model called Masked Visual Actions uses just 15 hours of robot video to predict action outcomes and generate robot behavior by representing motions as spatiotemporal pixel masks, achieving cross‑embodiment generalization, higher task success rates, and strong correlation between video evaluation and real‑world performance.

Video Predictioncross-embodiment generalizationinverse kinematics
0 likes · 9 min read
Masked Visual Actions: Controlling Robots with Only 15 Hours of Video
AIWalker
AIWalker
Mar 6, 2025 · Artificial Intelligence

How SCMHSA Improves Transformer Next‑Frame Prediction by Reducing Semantic Dilution

The paper introduces a Semantic‑Concentrated Multi‑Head Self‑Attention (SCMHSA) module and a new embedding‑space loss to address semantic dilution and loss‑target mismatch in Transformer‑based video next‑frame prediction, demonstrating significant PSNR and MSE gains across four benchmark datasets.

Embedding LossSCMHSASemantic Dilution
0 likes · 23 min read
How SCMHSA Improves Transformer Next‑Frame Prediction by Reducing Semantic Dilution