DeepMind’s GenCeption Shows Video Generation Can Serve as a General‑Purpose Vision Learner
DeepMind’s new GenCeption paper demonstrates that a pretrained text‑to‑video diffusion model can be transformed into a unified visual‑understanding system that handles depth, surface normal, segmentation, camera pose and 3D keypoint tasks, achieving performance comparable to specialist models while requiring dramatically fewer labeled examples.
