The purpose of this dataset is to pre- or post-train embedding models on text similarity tasks.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
The data generation process described in this paper was followed:
https://arxiv.org/pdf/2401.00368
Compute sponsored by Arrow Denmark and… See the full description on the dataset page:
https://huggingface.co/datasets/ThatsGroes/synthetic-from-unit-triple-tasks-norwegian.