Embedder is a multilingual triplet dataset designed for training and evaluating sentence embedding models using contrastive or triplet loss. It contains 1m examples across 11 Indic languages and English + 100extra langs, derived from the Samanantar parallel corpus and opus 100. Each example is structured as a triplet: (anchor, positive, negative).
This dataset is ideal for building… See the full description on the dataset page: https://huggingface.co/datasets/Parveshiiii/Embedder.