Tagged articles

I3D-ViT

2 articles · Page 1 of 1
Data Party THU
Data Party THU
Jul 26, 2026 · Artificial Intelligence

VideoChat3: Open-Source Full-Stack Video Understanding Model Linking Perception, Understanding, and Interaction

VideoChat3 is a 4‑billion‑parameter multimodal large model that unifies short‑video, long‑video, and streaming video understanding through a native spatiotemporal encoder, adaptive resolution budgeting, and four‑stage training, achieving competitive accuracy while dramatically reducing visual token count and inference cost.

Adaptive ResolutionI3D-ViTMultimodal LLM
0 likes · 15 min read
VideoChat3: Open-Source Full-Stack Video Understanding Model Linking Perception, Understanding, and Interaction
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Open VideoChat3: A Full‑Stack 4B Video Understanding Model for Short, Long, and Streaming Videos

VideoChat3 is a 4‑billion‑parameter multimodal model that introduces an Inflated 3D Vision Transformer and an adaptive frame‑resolution mechanism to efficiently handle short clips, hour‑long videos, and live streams, achieving competitive benchmarks while being fully open‑source.

Adaptive Frame ResolutionI3D-ViTMultimodal LLM
0 likes · 13 min read
Open VideoChat3: A Full‑Stack 4B Video Understanding Model for Short, Long, and Streaming Videos