This dataset is designed for multilingual information retrieval training.
Compared with raw Wikipedia paragraph dumps, it provides cleaner and more practical supervision by pairing title/section-driven queries with relevant paragraph-level documents and applying rule-based filtering to remove low-value sections and noisy fragments.
It is intended for IR model training, including contrastive learning and retrieval/reranking objectives.… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-ir-pairs.