This dataset is designed for training and evaluating Vision-Language Models (VLMs) on physical property estimation tasks. It contains 216 high-quality, balanced annotations mapping visual instances to their ground-truth physical properties (mass and friction).
Total Samples: 216 (Cleaned & Balanced)
Source Videos: 5 YouTube kitchen/cooking videos
Unique Object Instances: 162
Unique GT Object IDs: 78
Category… See the full description on the dataset page:
https://huggingface.co/datasets/sanskxr02/physvlm-dataset.