BioMedFlickr is a biomedical image–caption retrieval benchmark built from
public Flickr pathology / microscopy albums. Each example is a single image
paired with a cleaned English caption. The Hugging Face split matches the
evaluation set used in the EVVLM retrieval notebook
(dev_retrival-Copy1.ipynb): captions shorter than 10 characters after
cleaning are dropped, leaving 7185 test pairs.
Images are stored as original JPEGs inside parquet files so the dataset… See the full description on the dataset page:
https://huggingface.co/datasets/Alejandro98/BioMedFlickr.