NumerosityVLM is a controlled synthetic benchmark for evaluating numerosity perception in vision-language models (VLMs). It disentangles numerosity from correlated visual factors by systematically varying object size, spatial arrangement, and appearance cues.
The benchmark contains 10,800 images across six controlled conditions, three object categories, and twelve numerosity levels spanning 1… See the full description on the dataset page:
https://huggingface.co/datasets/fuy3/NumerosityVLM.