VideoChat3: Open-Source Full-Stack Video Understanding Model Linking Perception, Understanding, and Interaction
VideoChat3 is a 4‑billion‑parameter multimodal large model that unifies short‑video, long‑video, and streaming video understanding through a native spatiotemporal encoder, adaptive resolution budgeting, and four‑stage training, achieving competitive accuracy while dramatically reducing visual token count and inference cost.
