Research · Note · October 2026

How to evaluate a generated 3D object.

A generated 3D object is compared with a reference kept out of training: a held-out object. Each metric answers one question about the gap between the two shapes, and each comes with conventions that change its value.

Chamfer distance

Sample points on both surfaces. For each point of the generated shape, find the nearest point of the reference, and do the same in the other direction. The Chamfer distance averages these nearest-neighbour distances. Lower is better.

CD(A, B) = meana ∈ A minb ∈ B ‖a − b‖ + meanb ∈ B mina ∈ A ‖b − a‖

The conventions vary from paper to paper: squared or plain distances, a sum or a mean of the two directions, the number of sampled points, and the scale the shapes are normalised to (a unit cube, a unit sphere or real dimensions). Two papers can report different Chamfer distances for the same pair of shapes, so a value is comparable when these choices are stated next to it.

F-score at a threshold

Choose a distance threshold, for example 1% of the object's size. Precision is the share of generated points within that distance of the reference. Recall is the share of reference points within that distance of the generated shape. The F-score is their harmonic mean, from 0 to 1, and higher is better.

Tatarchenko and colleagues (CVPR 2019) recommended it for 3D reconstruction because it reports how much of the surface is right, where a Chamfer average can be pulled up by a few distant points.

Volumetric IoU

Voxelise both shapes, or sample occupancy on a grid, then divide the volume they share by the volume either one covers. IoU expects closed, watertight shapes, and it weighs large volumes more than thin details such as chair legs.

Normal consistency

Compare the orientation of the two surfaces at matching points: the mean absolute cosine between their normals, from 0 to 1. It separates a shape that sits in the right place with a smooth surface from one that sits there with the wrong surface detail.

Measures for generated objects

When a prompt allows many valid shapes, papers also render the objects and compare images, with FID or KID, measure agreement between the prompt and the renders with a CLIP score, and run studies where people choose between two results.

An editable asset adds its own checks: the error on stated dimensions in centimetres, the part count compared with the reference, and whether each part moves on its own. Our note on editable 3D generation describes these properties.

What a result should state

Sandflow is a frontier lab building foundation models for 3D, and we report our results this way. Our first model is on its way.

Related notes