Data Party THU
Sep 7, 2026 · Artificial Intelligence
Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
This survey reviews vision-language models for egocentric video, analyzing why standard VLMs fail on first-person footage and how hand-object interaction, spatiotemporal reasoning, graph-based representations, and semantic alignment bridge the gap toward embodied AI applications.
Egocentric VideoEmbodied AIGraph Neural Networks
0 likes · 23 min read
