Project Page | Paper | GitHub
VisGym consists of 17 diverse, long-horizon environments designed to systematically evaluate, diagnose, and train Vision-Language Models (VLMs) on visually interactive tasks. In these environments, agents must select actions conditioned on both their past actions and observation history, challenging their ability to handle complex, multimodal sequences.
This dataset contains trajectories and interaction data… See the full description on the dataset page:
https://huggingface.co/datasets/VisGym/visgym_data.