This dataset contains multilingual pairs of related Wikipedia paragraphs.
Each pair is sampled from the same article and the same section, so the two paragraphs are topically related but not necessarily paraphrases.
The dataset is intended for large-scale contrastive learning, representation learning, and retrieval-style training where broad positives are useful.