CC-Bench: A Cognitive Conflict Benchmark for MLLMs in Safety-Critical Visual Inspection
CC-Bench is a joint medical-industrial benchmark for evaluating whether multimodal large language models (MLLMs) remain visually grounded when plausible textual context conflicts with image evidence. The benchmark reorganizes public anomaly datasets into a unified four-way multiple-choice QA format for high-risk visual inspection.
This repository currently contains: