This dataset contains a set of candidate documents for second-stage re-ranking on webis-touche2020
(test split in BEIR). Those candidate documents are composed of hard negatives mined from
gtr-t5-xl as Stage 1 ranker
and ground-truth documents that are known to be relevant to the query. This is a release from our paper
Policy-Gradient Training of Language Models for Ranking, so
please cite it if using this dataset.
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/NeuralPGRank/webis-touche2020-hard-negatives.