PubChem & Wikipedia English-Hindi Paragraph Pair Classification
This dataset is a multilingual extension of the PubChem & Wikipedia Paragraphs Pair Classification dataset. It includes pairs of paragraphs in English and Hindi (sent1 and sent2) with a binary labels column indicating whether the paragraphs describe the same entity (1) or different entities (0).