

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I build an AR CAD app for iOS (http://graphite3d.com), and the gap I keep hitting is that world models are great at reconstructing space but bad at editable geometry — users want parametric surfaces they can dimension, not meshes. Curious whether Atlas exposes structured scene primitives.by iangraphite
- This seems like by far the best model yet for reconstructing 3D spaces from sparse images. It looks like you could reconstruct your whole house with pretty good fidelity from a dozen or so images taken on your phone.
They show it working with videos that have motion, but it seems like time is always frozen while the camera is moving, and they always return to a ground truth camera view before advancing time again. Maybe the temporal consistency isn't very good? This surprises me given how well it understands space. I guess modeling physics and time is the next step in the development of this kind of model.
by modeless - These models will open so many new possibilities for 3d workflows. Reminds me of a few years ago when I wad beta testing Stable Diffusion and realising it was going to change the way people design, visualize and present things forever. These models feel similar (if they work as advertised). I tested the first version of the World Labs model and it was OK but quite limited (quality of reconstruction, type of inputs and outputs) this one feels like a big breakthrough happened, and the model is probably much bigger. Looking forward to test it.by pablonaj
- It's a promising approach - and the demo goes to show just how advanced and robust "3D from 2D" reconstruction is now.
Dedicated depth sensors used to be a must on advanced robotics platforms - the only way to get anything close to reliable 3D point clouds was to spin a LiDAR. But by now, I wouldn't be surprised to see more and more robots ship with smartphone-like camera blocks - varying FoVs and focal depths, but not a lot of explicit depth sensing, if any at all.
Also, I wonder if this very model can be retrofit into a true robotics VLA? If it already takes text and image guidance, performs autoregressive diffusion of novel views, and handles temporal dynamics - why not diffusion of actions too?
by ACCount37 - What exactly does world model mean? Ive seen it used so many times in so many ways to just describe SOTA anything its lost its meaning.by thinkingkong
- I'm a cofounder at World Labs - happy to answer questions about Atlas!by jcjohns
- This is incredible. One potential application that I'm thinking about already is the rapid iteration of video-game map blocking. Being able to drop in some 'initial state' configuration and then have it procedurally generate a handful of alternative configurations could make rapid prototyping a significantly quicker experience, especially if you wanted to see what a potential end result could look like.
Furthermore, being able to extract and process world geometry and 3D objects from Atlas could reduce friction in the early stages of indy development, where developer time is stretched thinner.
I'm very excited about AI tooling moving forward if this is a glimpse into the future.
by Vakaiser - The blog post doesn't seem to mention what strikes me as the most interesting application of a model like this, namely extracting semantic information from its latent space. It mentions robotics applications, but only in the context of generating realistic world models for simulation.
If you have a robot deployed in an environment, generating synthetic views of the environment you're in doesn't have any obvious value. What does have obvious value is the latent knowledge that the model could have used to generate those synthetic views.
For instance, the fact that Atlas is capable of identifying regions of the input images that look like "floors", and smoothly interpolating them and filling in gaps with more floor, suggests that it has a concept of "floor-like walkability" which it's learned from the examples in its training data. And being able to identify the regions of 3D space that correspond to that semantic label would obviously be useful for robot path planning.
There's plenty of literature about e.g. using neural networks to estimate walkable areas from a point cloud. And you could imagine just bolting one of those methods to the front of Atlas, using the synthesized point cloud (instead of traditional photogrammetry or LIDAR) as input. But that seems like it's throwing away a lot of potentially useful semantic information, on top of being needlessly inefficient.
by teraflop