AugQ-CC is an unsupervised augmented dataset for training retrievers used in AugTriever: Unsupervised Dense Retrieval by Scalable Data Augmentation.
It consists of 52.4M pseudo query-document pairs based on Pile-CommonCrawl.
@article{meng2022augtriever,
title={AugTriever: Unsupervised Dense Retrieval by Scalable Data
Augmentation},
author={Meng, Rui and Liu, Ye and Yavuz, Semih and Agarwal, Divyansh and Tu, Lifu and Yu, Ning and Zhang, Jianguo and Bhat, Meghana and Zhou, Yingbo}… See the full description on the dataset page:
https://huggingface.co/datasets/memray/AugTriever-AugQ-CC.