CLIRudit is a dataset for academic Cross-lingual information retrieval (CLIR), consisting of English queries and French documents, based on Érudit, a non-profit publishing platform based in Quebec, Canada.
The CLIRudit dataset follows a TREC-style structure with three main components:
Queries: Generated from English keywords of research articles by creating all possible three-keyword combinations.
For example, an article with keywords {A, B, C, D} would… See the full description on the dataset page:
https://huggingface.co/datasets/ftvalentini/clirudit.