World Labs, co-founded by AI researcher Fei-Fei Li, has announced Atlas, a world model that generates, reconstructs, and simulates 3D scenes from just a few images. The company claims it beats specialized models at their own tasks, which could make many of them unnecessary.
Since its founding, World Labs has pursued the goal of “spatial intelligence”, the idea that AI should understand 3D space the way humans do. Atlas is the company’s first model built to do that at scale. Rather than producing flat images or video clips, it grasps how a scene looks from any angle and how it changes over time.
World Labs describes Atlas as an omni-model trained from scratch on text, images, video, and 3D data. Every input gets anchored to a specific position in 3D space rather than processed as a flat sequence. The company calls this shared spatial understanding “spatial context,” and it’s what the model uses to generate each new frame or viewpoint. According to World Labs, this anchoring separates Atlas from pure language or video models.
Fei-Fei Li laid out this exact problem in a November 2025 essay. Current multimodal language models and video diffusion models break data into one- or two-dimensional sequences, she argued, which makes even simple spatial tasks needlessly hard. What’s needed are architectures that organize tokenization, context, and memory in a 3D- or 4D-aware way.
One minute of video at 1440p
For camera-controlled generation, Atlas takes one or more images and produces new views at freely chosen camera positions and angles. Camera movement is passed as a direct geometric input rather than described through text prompts, as many video models require.
The model outputs up to one minute of video at 1440p. Users can control every shot themselves instead of “pulling the lever on a slot machine,” as World Labs put it, drawing a line between controlled generation and random output.

For spatial reconstruction, Atlas rebuilds real scenes from as few as one to several dozen input images without special capture equipment. The more images it receives, the less it has to fill in from its own knowledge.
With just two or three images, Atlas delivers faithful results and outperforms specialized 3D models, according to World Labs. It can also handle over a hundred inputs. In one demo, the model progressively assembles Stanford’s Main Quad from two to 25 ground-level photos and generates aerial views far above the campus.

This is where existing models tend to fall apart. In a comparison within the OpenWorldLib framework, systems like VGGT and InfiniteVGGT showed geometric inconsistencies and blurry textures as soon as the camera moved significantly.
Native 3D output and robotics simulation
Atlas can output results as actual 3D data, not just images or video, because it processes depth information alongside RGB. Supported formats include point clouds and 3D Gaussian splats, which build a scene from many small spatial data points that can be viewed smoothly from any angle. This matches the representation used in Marble, the company’s existing product.

As a simulator, Atlas models space and time together. From footage captured by just a few cameras, it can produce a “bullet time” effect that freezes a scene and lets users view it from otherwise impossible angles. The demo footage was shot with a handful of smartphones and action cameras, not professional gear.
For robotics, Atlas serves as a real-to-sim tool. It reconstructs a room and generates the image and depth data that a simulated robot’s sensors would see along its path. From just a few photos, users can simulate and vary grasping and movement tasks by swapping out objects, positions, lighting, or backgrounds. The goal is to produce diverse training data for robots without capturing every situation in the real world.
World Labs showed this approach in August 2026 with its real-to-sim-to-real engine as a standalone product. That engine creates thousands of variants from a single real-world task and trains control models entirely in simulation. On five robot platforms, the models ran for an hour each without human intervention, according to the company. The technology came from SceniX, a startup World Labs acquired in July.
Text-to-image generation isn’t the main focus, the company says, but Atlas can also follow complex prompts, render text, produce different visual styles, and create 360-degree panoramas.
Speed from language models, quality from diffusion
Atlas combines ideas from both language models and video models. It generates output piece by piece like a language model, so it can use the same speedup techniques, such as KV caching. But it also uses the diffusion principle from image and video models, gradually filtering output out of noise. That side gives it access to methods that shorten the denoising process or boost image quality.
World Labs says no single benchmark captures what Atlas can do, but points to two sets of tests. In camera-controlled generation judged by external human evaluators, and in few-view 3D reconstruction, Atlas outperforms more specialized models.
Human evaluators preferred Atlas in 75 percent of comparisons against MiniMax H3, 81 percent against Gemini Omni Flash, 86 percent against Happy Horse 1.1, 93 percent against, and 94 percent against Seedance 2.5. For reconstruction, Atlas leads with a median error of 25.3, ahead of Pi3X and VGGT-Ω 1B.

The company says Atlas’s performance improves with more training compute and expects that trend to hold as it scales. Atlas will power future versions of Marble and other products, and is currently available through an early-access program for select partners.

From walkable photos to an omni-model
World Labs was founded in 2024 by Fei-Fei Li, who created ImageNet and led Google Cloud’s AI division from 2017 to 2018. The company at launch from Andreessen Horowitz, AMD, Intel, and Nvidia.
A first system in late 2024 turned, though users could only move a few virtual meters before hitting invisible boundaries. Marble followed in November 2025. In February 2026 came a $1 billion funding round from Autodesk, Andreessen Horowitz, Nvidia, and AMD. Bloomberg had previously reported talks at a $5 billion valuation.
What counts as a world model remains contested among researchers. An international team led by Peking University proposed a unified definition in April 2026 through OpenWorldLib, excluding pure text-to-video models because they lack feedback loops with the real world. 3D reconstruction and simulators like those in Atlas qualify as core building blocks in that framework because they provide environments where physical rules can be verified.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.
Read on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: “AI Radar” — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI








