Hard image-pair contrasts mined from
mtybilly/PubMedVision-Alignment-VQA,
the flat single-image medical VQA derived from the upstream
FreedomIntelligence/PubMedVision.
Each row is a pair of two medical images that are:
visually similar but not identical — same modality + body part bucket, BiomedCLIP image cosine ∈ [0.85, 0.99]
same-intent question — BiomedCLIP text-encoder cosine ≥ 0.73 (admits paraphrased templates: "describe /… See the full description on the dataset page:
https://huggingface.co/datasets/mtybilly/PubMedVision-Diff.