This dataset contains approximately 36 million Wikipedia passages from the "Multi-task retrieval for knowledge-intensive tasks" paper. It is also the retrieval corpus used in the paper Chain-of-Retrieval Augmented Generation.
id: A unique identifier for each passage.
title: The title of the Wikipedia page from which the passage originates.
contents: The textual content of the passage.
wikipedia_id: The unique identifier for the Wikipedia page, used for… See the full description on the dataset page:
https://huggingface.co/datasets/corag/kilt-corpus.