VS-Bench is a multimodal benchmark for evaluating VLMs in multi-agent environments. We evaluate fourteen state-of-the-art models in eight vision-grounded environments with two complementary dimensions, including offline evaluation of strategic reasoning by next-action prediction accuracy and online evaluation of decision-making by normalized episode return.
@article{xu2025vs,
title={VS-Bench: Evaluating VLMs for Strategic Reasoning… See the full description on the dataset page:
https://huggingface.co/datasets/zelaix/VS-Bench.