Tagged articles

Long Video

7 articles · Page 1 of 1
Data Party THU
Data Party THU
Jul 26, 2026 · Artificial Intelligence

VideoChat3: Open-Source Full-Stack Video Understanding Model Linking Perception, Understanding, and Interaction

VideoChat3 is a 4‑billion‑parameter multimodal large model that unifies short‑video, long‑video, and streaming video understanding through a native spatiotemporal encoder, adaptive resolution budgeting, and four‑stage training, achieving competitive accuracy while dramatically reducing visual token count and inference cost.

Adaptive ResolutionI3D-ViTLong Video
0 likes · 15 min read
VideoChat3: Open-Source Full-Stack Video Understanding Model Linking Perception, Understanding, and Interaction
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Open VideoChat3: A Full‑Stack 4B Video Understanding Model for Short, Long, and Streaming Videos

VideoChat3 is a 4‑billion‑parameter multimodal model that introduces an Inflated 3D Vision Transformer and an adaptive frame‑resolution mechanism to efficiently handle short clips, hour‑long videos, and live streams, achieving competitive benchmarks while being fully open‑source.

Adaptive Frame ResolutionI3D-ViTLong Video
0 likes · 13 min read
Open VideoChat3: A Full‑Stack 4B Video Understanding Model for Short, Long, and Streaming Videos
JD Tech Talk
JD Tech Talk
Jun 11, 2026 · Artificial Intelligence

How JD’s Open‑Source JoyAI‑Echo Overcomes the Three Biggest Long‑Video Generation Challenges

JoyAI‑Echo, JD’s newly open‑sourced long‑video generation framework, tackles character inconsistency, voice instability, and slow rendering by introducing a cross‑modal memory bank, memory‑driven training with DMD for 7.5× speedup, a conversational Director Agent, and real‑time super‑resolution, achieving leading benchmark scores and high user preference.

AI Video GenerationDirector AgentLong Video
0 likes · 6 min read
How JD’s Open‑Source JoyAI‑Echo Overcomes the Three Biggest Long‑Video Generation Challenges
Kuaishou Tech
Kuaishou Tech
Jun 11, 2026 · Artificial Intelligence

Keye-VL-2.0 Brings DeepSeek Sparse Attention to Multimodal AI – Report Released

Keye‑VL‑2.0, an open‑source MoE multimodal foundation model, tackles hour‑level video understanding and agentic intelligence by embedding DeepSeek Sparse Attention into a GQA‑based architecture, enabling near‑lossless 256 K token context, four‑stage pre‑training, diverse RL distillation techniques, and achieving state‑of‑the‑art results on long‑video benchmarks, with weights publicly released.

Long VideoMoEPretraining
0 likes · 8 min read
Keye-VL-2.0 Brings DeepSeek Sparse Attention to Multimodal AI – Report Released
Machine Heart
Machine Heart
Jun 6, 2026 · Artificial Intelligence

How JoyAI‑Echo Generates 5‑Minute AI Videos in One Shot and Ditches the Blind‑Box Approach

JoyAI‑Echo, an open‑source framework from JD, enables fully consistent five‑minute AI video generation with a single pass, offering non‑linear editing, high‑resolution output up to 1472×2560, and a suite of memory‑driven techniques that overcome the long‑video bottlenecks of earlier models.

AI Video GenerationDirector AgentJoyAI-Echo
0 likes · 12 min read
How JoyAI‑Echo Generates 5‑Minute AI Videos in One Shot and Ditches the Blind‑Box Approach
Machine Heart
Machine Heart
May 6, 2026 · Artificial Intelligence

Scal3R Enables Stable Kilometer-Scale 3D Reconstruction of Long Videos

Scal3R introduces test‑time training with a global‑context memory and synchronization mechanism that lets models train on and infer over ultra‑long video sequences, achieving accurate camera poses and dense point clouds for kilometer‑scale scenes while outperforming prior SLAM, SfM and streaming baselines on multiple benchmarks.

3D reconstructionLong VideoScal3R
0 likes · 11 min read
Scal3R Enables Stable Kilometer-Scale 3D Reconstruction of Long Videos
Kuaishou Tech
Kuaishou Tech
Aug 25, 2025 · Artificial Intelligence

How Context-as-Memory Enables Scene‑Consistent Long Video Generation

This article introduces the Context-as-Memory approach, which treats previously generated video frames as memory to achieve scene‑consistent interactive long video generation, and details a camera‑trajectory‑based memory retrieval mechanism that dramatically improves efficiency and performance over existing state‑of‑the‑art methods.

AILong VideoMemory Retrieval
0 likes · 7 min read
How Context-as-Memory Enables Scene‑Consistent Long Video Generation