This repository contains the generated pseudo queries used for training the DDRO (Direct Document Relevance Optimization) models.The pseudo queries were created following the DocTTTTTQuery approach to expand the training data for generative document retrieval.
pseudo_queries_msmarco.txt: Pseudo queries for the MS MARCO (MS300K) dataset.
pseudo_queries_nq.txt: Pseudo queries for the Natural Questions (NQ320K) dataset.
Each file maps… See the full description on the dataset page:
https://huggingface.co/datasets/kiyam/ddro-pseudo-queries.