Percept-V is a benchmark of 30 synthetic visual perception tasks, 200 samples each (6,000 samples total), designed to isolate perception from reasoning in vision-language models. Every task is procedurally generated from simple primitives — circles, lines, grids, shapes, colours — so a model that genuinely sees the image should solve it near-perfectly, and failures point at perceptual rather than reasoning limits.
Tasks span counting, colour identification, shape… See the full description on the dataset page:
https://huggingface.co/datasets/aggr8/Percept-V.