How to evaluate a generated 3D object.
A generated 3D object is compared with a reference kept out of training: a held-out object. Each metric answers one question about the gap between the two shapes, and each comes with conventions that change its value.
Chamfer distance
Sample points on both surfaces. For each point of the generated shape, find the nearest point of the reference, and do the same in the other direction. The Chamfer distance averages these nearest-neighbour distances. Lower is better.
CD(A, B) = meana ∈ A minb ∈ B ‖a − b‖ + meanb ∈ B mina ∈ A ‖b − a‖
The conventions vary from paper to paper: squared or plain distances, a sum or a mean of the two directions, the number of sampled points, and the scale the shapes are normalised to (a unit cube, a unit sphere or real dimensions). Two papers can report different Chamfer distances for the same pair of shapes, so a value is comparable when these choices are stated next to it.
F-score at a threshold
Choose a distance threshold, for example 1% of the object's size. Precision is the share of generated points within that distance of the reference. Recall is the share of reference points within that distance of the generated shape. The F-score is their harmonic mean, from 0 to 1, and higher is better.
Tatarchenko and colleagues (CVPR 2019) recommended it for 3D reconstruction because it reports how much of the surface is right, where a Chamfer average can be pulled up by a few distant points.
Volumetric IoU
Voxelise both shapes, or sample occupancy on a grid, then divide the volume they share by the volume either one covers. IoU expects closed, watertight shapes, and it weighs large volumes more than thin details such as chair legs.
Normal consistency
Compare the orientation of the two surfaces at matching points: the mean absolute cosine between their normals, from 0 to 1. It separates a shape that sits in the right place with a smooth surface from one that sits there with the wrong surface detail.
Measures for generated objects
When a prompt allows many valid shapes, papers also render the objects and compare images, with FID or KID, measure agreement between the prompt and the renders with a CLIP score, and run studies where people choose between two results.
An editable asset adds its own checks: the error on stated dimensions in centimetres, the part count compared with the reference, and whether each part moves on its own. Our note on editable 3D generation describes these properties.
What a result should state
- The metric and its formula, including squared or plain distances.
- The number of sampled points and how they were sampled.
- The normalisation of the shapes, or their real dimensions.
- The threshold used for the F-score.
- The held-out set: its categories and the number of objects.
- The hardware and the time per object, when speed is part of the claim.
Sandflow is a frontier lab building foundation models for 3D, and we report our results this way. Our first model is on its way.