Paper: Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
WM-ABench is a comprehensive benchmark that evaluates whether Vision-Language Models (VLMs) can truly understand and simulate physical world dynamics, or if they rely on shortcuts and pattern-matching. The benchmark covers 23 dimensions of world modeling across 6 physics simulators with over 100,000… See the full description on the dataset page:
https://huggingface.co/datasets/maitrix-org/WM-ABench.