Repository |
Paper |
ArXiv
RLHN is a cascading LLM framework designed to accurately relabel hard negatives in existing IR/RAG training datasets, such as MS MARCO and HotpotQA.
This Tevatron dataset (680K training pairs) contains the original queries, positives and hard negatives after dropping each training pair with a single false negative.
This repository contains the training pairs that can be used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/rlhn/remove-680K.