This is the synthetic dataset used for training Dutch embedding models as described in MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch.
Each sample contains the following fields:
task_type: Type of the embedding task; one of the five categories:
sl (short-long): retrieval
ls (long-short): classification
ss (short-short): clustering
ll (long-long): clustering
sts (semantic text similarity): semantic text… See the full description on the dataset page:
https://huggingface.co/datasets/clips/SynEmbedNL.