A large-scale 3D vision-language question-answering dataset for point-grounded,
metric-aware 3D scene understanding. Built on ScanNet and ScanNet++
with dense object and part annotations, the dataset spans eight downstream
reasoning tasks at both object and part granularity, plus multi-turn
dialogue that composes them.
Answers are both:
point-grounded: explicitly tied to the referred 3D region via
label
markup, and
metric: physical quantities (size… See the full description on the dataset page: https://huggingface.co/datasets/amolharsh/Ground3D_Dataset.