Shot composition stops being a decision made before the camera rolls: the scene gets filmed first, the angle and framing settled afterward.
World Labs, the AI world-model company led by Fei-Fei Li, announced Atlas. It takes ordinary photographs or phone frames — as few as one, or up to several dozen — works out a three-dimensional version of the space in them, and then generates the footage a camera would have captured on a path nobody walked, up to a minute of it at 1440p.
In a demonstration shot on a phone, Atlas stops the motion mid-scene, moves a virtual camera around the environment it has reconstructed, and lets the footage run on from a different viewpoint. Video is not the only thing that comes out: the model also emits point clouds, Gaussian splats and depth maps, and — for robotics — the colour and depth readings a robot’s cameras would return along a given trajectory, which is training data for a machine that never entered the room.
The claim worth weighing is that a general model beats the specialists. World Labs reports a sparse-view reconstruction error of 25.3 against 28.7 for the nearest open-source system built for that job alone, and human preference of 75% to 94% on camera control, depending on the competitor. Every one of those numbers is the company’s own, and commercial systems were left out of the comparison.
Nobody outside can check them yet. Atlas is entering early access through a request form; World Labs has named no partners, published no price, set no date for general availability and released no weights. It says the model will power future versions of Marble, the product it already sells.