This repository contains the ScanNet dataset (3D scene data and 2D frame data) and refined annotations used for the paper Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding.
TAB is a dynamic agentic framework designed for zero-shot 3D Visual Grounding (3D-VG). By operating directly on raw RGB-D streams, TAB reformulates 3D… See the full description on the dataset page:
https://huggingface.co/datasets/AntonioJun/indoor-3d-bbox-indices.