MedRank-DecisionGrade is a medical LLM pairwise preference evaluation dataset for studying harm-aware and annotator-aware ranking.
questions.jsonl: 400 public medical QA questions.
generations.jsonl: 2,000 model generations from five open-weight LLMs.
pairs.jsonl: 2,000 pairwise comparisons.
annotations_all_validated.jsonl: 1,950 human annotation records from two trained annotators and one physician, covering 1,500 unique pair IDs.… See the full description on the dataset page:
https://huggingface.co/datasets/medrank-benchmark/medrank-decisiongrade.