This repository hosts the ReMedQA dataset, introduced in our paper: "ReMedQA: Are We Done With Medical Multiple-Choice Benchmarks?"
While medical multiple-choice question answering (MCQA) benchmarks often report near-human accuracy, raw accuracy alone does not reliably measure a model’s true competence. Models may change answers under minor perturbations, exposing fragility and lack of robustness.
ReMedQA addresses this limitation by extending standard… See the full description on the dataset page:
https://huggingface.co/datasets/disi-unibo-nlp/ReMedQA.