We release the raw passages, embeddings, and index of MassiveDS.
Website:
https://retrievalscaling.github.io
We release two versions of MassiveDS:
MassiveDS-1.4T, which contains 1.4T tokens in the datastore.
MassiveDS-140B, which is a subsampled version containing 140B tokens in the datastore.
raw_data: plain data in JSONL files.
passages: chunked raw passages with passage IDs. Each passage is chunked to have no more than 256 words.
embeddings: embeddings of the passages… See the full description on the dataset page:
https://huggingface.co/datasets/rulins/MassiveDS-140B.