A medical-domain benchmark dataset for evaluating retrieval, reranking, and RAG systems under low lexical overlap and high semantic difficulty.
⚠️ Designed to penalize shallow matching.High scores from lexical methods (e.g., BM25) may indicate shortcut exploitation, not real understanding.
reduce lexical similarity between queries and relevant content
increase semantic diversity across… See the full description on the dataset page:
https://huggingface.co/datasets/ReDiX/ReDiX-Benchmark-Medical-ita.