This HF dataset contains the snippets from the PubMed corpus used in MedRAG. It can be used for medical Retrieval-Augmented Generation (RAG).
News
(02/26/2024) The "id" column has been reformatted. A new "PMID" column is added.
Dataset Details
Dataset Descriptions
PubMed is the most widely used literature resource, containing over 36 million biomedical articles.
For MedRAG, we use a PubMed subset of 23.9 million… See the full description on the dataset page: https://huggingface.co/datasets/minsu/medrag_pubmed.