Artifacts for Geometric Code, a training-free perception-to-geometry pipeline
that converts monocular RGB video into a structured, metric spatial code and
supplies it to vision-language models (VLMs) as additional context for spatial
reasoning.
A perception layer segments and classifies objects (SAM 3) and recovers per-frame
metric depth, camera pose, and intrinsics (Depth Anything 3). A deterministic
geometric… See the full description on the dataset page:
https://huggingface.co/datasets/SpatialCode/Geometric-Code.