VLbenchy is a 500-sample procedural vision-language benchmark for evaluating the fundamental visual understanding capabilities of VL models across 10 distinct task types.
Each sample contains a procedurally generated 336×336 PNG image paired with a natural-language question, a ground-truth answer, and multiple-choice options. All images are unique — generated with randomised shapes, colors, sizes, and layouts.
id… See the full description on the dataset page:
https://huggingface.co/datasets/Tentlow/VLbenchy.