This dataset contains multilingual Wikipedia lead-text pairs designed for large-scale representation learning, contrastive pretraining, clustering, and classification-oriented embedding training.
Each example pairs the lead text of one Wikipedia article with the lead text of another article that is judged to be strongly related in a broad topical sense.
The core signal comes from Wikipedia hyperlinks, especially links attached to lead… See the full description on the dataset page:
https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-ir-related-lead-pairs.