This HF dataset contains the chunked snippets from the Wikipedia corpus used in MedRAG. It can be used for medical Retrieval-Augmented Generation (RAG).
News
(02/26/2024) The "id" column has been reformatted. A new "wiki_id" column is added.
Dataset Details
Dataset Descriptions
As a large-scale open-source encyclopedia, Wikipedia is frequently used as a corpus in information retrieval tasks.
We select Wikipedia as one… See the full description on the dataset page: https://huggingface.co/datasets/minsu/medrag_wikipedia.