The purpose of this dataset is to pre- or post-train embedding models for Danish on text similarity tasks.
The dataset is structured for training using InfoNCE loss (also known as SimCSE loss, Cross-Entropy Loss with in-batch negatives, or simply in-batch negatives loss), with hard-negative samples for the tasks of retrieval and unit-triplet. Beware that if fine-tuning the unit-triplets for… See the full description on the dataset page:
https://huggingface.co/datasets/DDSC/nordic-embedding-training-data.