World Labs announced Atlas on September 1, 2026, calling it "the world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D." The pitch, per the company: "model the world, move the camera, and simulate space & time."
What Atlas actually does
Atlas is a multimodal autoregressive diffusion transformer — one model handling text, images, video, and 3D data together inside what World Labs calls a unified spatial framework, rather than routing between separate specialist models. In practice that means it can generate image and video frames at up to 1440p and one minute in length with precise, controllable camera movement; reconstruct a full 3D scene from as little as a single input image; reframe existing video to simulate different camera paths through the same space; and compose several posed images of a scene into one consistent 3D reconstruction. World Labs says Atlas beats existing open-source reconstruction models on this last capability, though that's a vendor claim rather than an independently verified benchmark result.
The generation and reconstruction halves being the same model, rather than a generator feeding a separate reconstruction pipeline, is the actual architectural claim here — everything else follows from treating "create a scene" and "reconstruct a scene" as two views of one underlying representation.
From Marble to Atlas
Atlas is the fourth step in a fairly deliberate sequence. Marble shipped in November 2025 as World Labs' first commercial product: 3D world generation from images or text, sold direct to users. In January 2026 the company opened the World API, giving developers and robotics teams programmatic access to that same generation capability instead of only a consumer app. In July 2026 it acquired SceniX, a robotics simulation company founded by Columbia professors Yunzhu Li and Changxi Zheng, built around a "real-to-sim-to-real" pipeline: digitize a real environment, train a robot's policy inside the digital twin, deploy the learned behavior back into the physical world. Atlas is the model that pipeline needs — reconstructing accurate 3D scenes from real footage is the "real-to-sim" half of that loop, and it's exactly the capability World Labs is now emphasizing over Marble's original consumer framing.
Early access is rolling out over the coming weeks rather than shipping broadly on day one.
The billion-dollar thesis behind it
World Labs raised $230 million to launch in September 2024 and closed a $1 billion round in February 2026 at roughly a $5 billion valuation, with AMD, Nvidia, Autodesk, Emerson Collective, Fidelity, and Sea among the investors — Autodesk alone put in $200 million and takes an advisory role. The thesis underneath all of it, in Fei-Fei Li's own framing, is that world models rather than language models are AI's next frontier — that teaching a system to reason about 3D space is a different and, in her view, more consequential problem than scaling text prediction further. Atlas is the clearest technical expression of that bet so far: not a chatbot with a 3D plugin, but a model built from the ground up to treat generation and spatial reconstruction as the same task.
Where this sits among world models
Atlas isn't the only thing called a "world model" this year. Google DeepMind's Genie 3, released in August 2025, generates real-time interactive environments at 720p and 24fps that a user can navigate frame by frame, with roughly a minute of memory. The two aren't really answering the same question. Genie 3 is optimized for live, playable interactivity; Atlas is optimized for precise, camera-controlled output plus explicit 3D reconstruction from minimal input — closer to a production and simulation tool than an interactive world to walk around in. Both companies are converging on "spatial intelligence" as the frame, but from opposite ends of the same problem.