Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
This survey reviews vision-language models for egocentric video, analyzing why standard VLMs fail on first-person footage and how hand-object interaction, spatiotemporal reasoning, graph-based representations, and semantic alignment bridge the gap toward embodied AI applications.
