SuanNi
Jul 15, 2026 · Artificial Intelligence
DeepMind’s Video Generation Model Becomes a General Visual Intelligence – He Kaiming’s Involvement
GenCeption repurposes a 140‑billion‑parameter text‑to‑video diffusion model into a single‑step feed‑forward visual system that handles depth, segmentation, pose and other tasks via text prompts, achieves state‑of‑the‑art results with far fewer training frames, and demonstrates strong out‑of‑domain generalisation using synthetic data.
GenCeptionmultitask visionsynthetic data
0 likes · 10 min read
