A diagnostic benchmark that measures whether a vision-language scoring
interface executes Boolean operators (negation, conjunction, disjunction,
exclusion, NOR) or merely tracks which concepts are mentioned. It contains
1,695 image/caption-pair samples built from COCO val2017, all OWL-ViT-validated.
Introduced in Similarity Is Not Logic: Factored Inference for Dual-Encoder
Vision-Language Models (ICML 2026).