This dataset is designed to evaluate the Theory of Mind (ToM) capability of vision-language models (VLMs). It includes data collected from two simulators: ThreeDWorld and VirtualHome.
Data Sources
1. ThreeDWorld
In ThreeDWorld, all data are collected through manual keyboard control.Each sample in this simulator contains: