PIRE stands for Pdf Information Retrieval Evaluation.
This dataset is composed of manually created queries and a corpus of PDF files. For each query, the relevant documents, pages, and text passages relevant to the queries have been labeled.
This dataset can be used for evaluation of information retrieval strategies on PDF files. It was initially created in order to compare various PDF parsing and chunking tools for information… See the full description on the dataset page:
https://huggingface.co/datasets/Wikit/PIRE.