This dataset contains the dataset for NeurIPS D&B submission: Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Precise Visual Grounding. It includes all the data for S1, S2 and S3 defined in the paper:
260+ source images
550 + questions
Human annotated masks (for S1 ans S2)
class information and subclass information
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pio-benchmark/PIO.