A benchmark for visually conditioned action on complete multi-object 3D scenes, under a unified agent–environment loop. This repository hosts the dataset; the evaluation harness lives at github.com/Feinaldo2/SceneActBench.
Vision-language model (VLM) agents increasingly use tools to act on 3D
scenes rather than only describe them. Existing 3D benchmarks score textual
responses or single-object… See the full description on the dataset page:
https://huggingface.co/datasets/FEInaldo/SceneActBench.