World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial ReasoningWanyue Zhang et al., 2026
π MotivationVision-Language Models (VLMs) excel at static visual understanding but struggle with dynamic spatial reasoning, such as predicting how a sceneβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/WanyueZhang/World2VLM.