Fei‑Fei Li Unveils Atlas: First Multimodal World Model to Reconstruct 3D Worlds from One Image

Atlas, a multimodal autoregressive diffusion Transformer introduced by Fei‑Fei Li's World Labs, can generate pixel‑level camera‑controlled images and videos, perform high‑quality 3D reconstruction from a few photos, simulate space‑time for robotics, and outperforms state‑of‑the‑art models in quantitative evaluations.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Fei‑Fei Li Unveils Atlas: First Multimodal World Model to Reconstruct 3D Worlds from One Image

Fei‑Fei Li’s World Labs announced Atlas, a next‑generation multimodal world model that integrates a multimodal autoregressive diffusion Transformer (an “omni model”) to handle text, images, camera poses, 3D depth maps and video as sequences anchored in 3‑D space.

Atlas illustration
Atlas illustration
Model the world, move the camera, simulate space and time.

Atlas can generate pixel‑level camera‑controlled images and videos from a single input image, outputting up to one‑minute 1440p video. It can reconstruct real‑world scenes from a few images, producing novel viewpoints and explicit 3‑D representations that surpass specialized 3‑D reconstruction models. By modeling space‑time jointly, Atlas can alter the viewpoint of existing video, create dramatic visual effects, and support Real‑to‑Sim workflows for robotics.

For robot simulation, feeding Atlas a handful of photos yields realistic RGB and depth data, enabling robots to train and test in diverse simulated environments.

Technically, Atlas combines a Rectified Flow diffusion backbone with a Transformer decoder. The diffusion process progressively denoises to generate high‑dimensional continuous data, while the Transformer provides a robust architecture for world‑modeling tasks. The model treats every task as a sequence‑to‑sequence problem: inputs form the prefix of the sequence, and outputs are generated autoregressively conditioned on that prefix.

Atlas leverages large‑language‑model techniques such as KV‑Cache, cache‑aware routing, and decoupled services, as well as modern diffusion tricks like distillation, classifier‑free guidance, and shifted‑noise schedules, together with an advanced VAE.

Quantitative evaluations on two key tasks—camera‑controlled generation and sparse‑view 3‑D reconstruction—show that Atlas outperforms state‑of‑the‑art video models, with larger gains on more complex camera trajectories, and achieves lower reconstruction error than the best open‑source 3‑D reconstruction baselines.

The team also reports early scaling evidence: performance continues to improve as model size grows, and the architecture is designed for future scaling.

Atlas is positioned as a milestone for embodied intelligence. Following Fei‑Fei Li’s earlier definition of world models as renderers, simulators, and planners, Atlas emphasizes the simulator role, bridging real photos/video, 3‑D space, robot sensor views, and controllable simulated environments. World Labs plans to integrate Atlas as the foundation for upcoming products such as Marble.

Atlas opens doors for many applications from visual effects to robotics.

Atlas is currently available for early access to select partners, with a public sign‑up link provided.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Roboticsmultimodal3D reconstructiondiffusion transformerworld modelAtlas
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.