Mechanical Reasoning is a multimodal benchmark designed to evaluate mechanical and physical reasoning in large vision-language models (VLMs). This dataset is introduced in the paper:
It is part of the Grow AI Like a Child benchmark initiative, which explores structured, curriculum-aligned evaluation for embodied intelligence.