This workspace evaluates VSI-Bench question answering with several input regimes:
raw video frames, perceived spatial codes from SAM3 + Depth Anything 3 caches,
ground-truth spatial codes from dataset annotations, and a deterministic symbolic
solver. The code is organized so important outputs are reproducible from fixed
inputs, fixed packages, fixed model checkpoints, and fixed SAM3/DA3 caches.
The repository intentionally separates three… See the full description on the dataset page:
https://huggingface.co/datasets/AntonioJun/workspace.