This is a subset of the MS MARCO v1.1 dataset by Microsoft, sampled for lightweight experimentation.
Original dataset: microsoft/ms_marco (v1.1)
Original paper: MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
Original authors: Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, Li Deng (Microsoft)
Randomly sampled 10,000 examples from the train split… See the full description on the dataset page:
https://huggingface.co/datasets/Lala8383/ms-marco-qa-10k.