[!NOTE]
This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first.
This repository contains training files to train text embedding models, e.g. using sentence-transformers.
All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example.
The JSON objects can come… See the full description on the dataset page:
https://huggingface.co/datasets/sentence-transformers/embedding-training-data.