A dataset used to train the not-yet-published querygen-v4 model from Nixiesearch. The dataset is a combination of multiple open query-document datasets in a format for Causal LLM training.
MSMARCO: 532751 rows
HotpotQA: 170000 rows
NQ: 58554 rows
MIRACL en: 1193 rows
SQUAD: 85710 rows
TriviaQA: 60283 rows
The train split is 900000 rows, and test split is 8491.… See the full description on the dataset page:
https://huggingface.co/datasets/nixiesearch/querygen-data-v4.