This is the dataset associated with the paper "BordIRlines: A Dataset for Evaluating Cross-lingual Retrieval-Augmented Generation" (link).
Code:
https://github.com/manestay/bordIRlines
The BordIRLines Dataset is an information retrieval (IR) dataset constructed from various language corpora. It contains queries and corresponding ranked docs along with their relevance scores. The dataset includes multiple languages, including English… See the full description on the dataset page:
https://huggingface.co/datasets/borderlines/bordirlines.