This dataset is designed for information retrieval (IR) research on scientific papers.
It contains a large corpus of papers collected from arXiv between 2015 and 2025, along with a set of synthetic queries and relevance judgments (qrels).
The dataset can be used to evaluate retrieval systems such as:
semantic search
document retrieval
neural… See the full description on the dataset page:
https://huggingface.co/datasets/c1khoa/papers_retrieval_arxiv.