SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) —
such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP — encode the spatial relations
between objects in a scene, rather than only their semantic identity.
📄 Paper: arXiv:2601.11729
💻 Code: github.com/gmum/SpaRRTa
🧱 Real-world (lego) split: turhancan97/SpaRRTa-Lego
🔬 Attention-analysis split (images +… See the full description on the dataset page:
https://huggingface.co/datasets/turhancan97/SpaRRTa.