What is a foundation model for 3D?
A foundation model is a large model trained on broad data, then adapted to many tasks. Researchers at Stanford named the category in 2021, with BERT and GPT-3 among their examples. Language, images and video each have theirs.
A foundation model for 3D, or 3D foundation model, does the same work for objects and scenes. It learns shape, structure and scale from a large body of 3D data, then serves many 3D tasks: building an object from a text or an image, completing a scan, varying a design.
What it outputs
3D has several representations, and each model works in one of them.
- Meshes: vertices and faces, the format of games, film and product visualisation.
- Point clouds: points sampled on a surface, the format of scanners and lidar.
- Gaussian splats: soft primitives fitted to photographs, fast to render.
- Implicit fields: a function of space, such as a signed distance field, that describes the surface.
- Programs: code that builds the object step by step, in a CAD kernel or in a 3D tool such as Blender.
The representation decides what a person can do with the result: render it, measure it, edit it or simulate it.
Why 3D needs its own models
An image shows an object from one viewpoint. A simulator, a game engine, a product configurator or a robot needs the object itself: its parts, its dimensions and its position in space. We think the next image, video and world models will rely on explicit 3D of this kind underneath.
What we look for in one
- Real-world scale: dimensions in metres, so an object fits its scene.
- Parts: each component kept as its own object, with a clear name and a pivot.
- Editable geometry: a result an artist can open in Blender and keep working on.
- Measured results: scores on held-out objects, with the metric stated, such as Chamfer distance or F-score.
Sandflow is a frontier lab building foundation models for 3D. Our first model is on its way.
Next note: what makes a generated 3D object editable.