This dataset is designed for training an embedding model using triplet loss. It contains triplets of affiliation strings (anchor, positive, negative) structured to teach a model to recognize when two strings refer to the same institution.
The dataset is sorted from easiest to hardest to facilitate curriculum learning, allowing the model to learn from simple examples before progressing to more challenging ones.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/triplet-loss-for-embedding-affiliations-sample-1.