The dataset CapRetrieval introduced in the EMNLP 2025 Finding paper: [Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings].
CapRetrieval is in Chinese; the according English version is available at CapRetrievalEn, sharing the same queries, passages and labels.
CapRetrieval evaluates the fine-grained embedding matching (dense passage retrieval), tailored towards a practical image search scenario:
Candidate passages are image… See the full description on the dataset page:
https://huggingface.co/datasets/lxucs/CapRetrieval.