PERCEIVE (Psychophysics-driven Elicitation for Routing Cost-Efficiency In
Vision-Language Evaluation) is a 4,801-sample document-image QA benchmark for
cost-aware VLM routing. Each sample carries psychophysical complexity annotations
(Visual Dependency Score, Reasoning Depth Score, Spatial Extent Score) and a
routing label identifying the cheapest model-budget configuration that answers
it correctly.
Routing labels are derived via a QUEST-style adaptive cascade achieving… See the full description on the dataset page:
https://huggingface.co/datasets/quantiphi-routing/perceive-benchmark.