A benchmark for evaluating whether LLMs understand hard physical constraints (walls, obstacles, grid boundaries) that require embodied spatial understanding.
Each sample is a 5×5 grid navigation problem with A*-computed ground truth and physics-engine-validated distractors.
1,000 samples | 4 choices each | 0% data contamination
Compatible with EleutherAI lm-evaluation-harness