MIRIAD is a curated million scale Medical Instruction and RetrIeval Dataset. It contains 5.8 million medical question-answer pairs, distilled from peer-reviewed biomedical literature using LLMs. MIRIAD provides structured, high-quality QA pairs, enabling diverse downstream tasks like RAG, medical retrieval, hallucination detection, and instruction tuning.
The dataset was introduced in our arXiv preprint.
from datasets… See the full description on the dataset page:
https://huggingface.co/datasets/MONARCH4842/miriad-5.8M.