This is a machine translation of the MS MARCO dataset.
This dataset can be used to train sentence embedding models.
In contrast to our previous translation,
an LLM (GPT-4o mini) was used for the translation.
This results in generally higher translation quality.
The dataset is based on the MS MARCO dataset.
We used a deployment of GPT-4o mini using the Microsoft Azure OpenAI APIs.