This is the original and unchanged german translated dataset (train split only) in original order from jphme/slimorca_dedup_german_experimental with added cosine-similarity scores. As no license was given for this version, I chose the MIT license from the original Open-Orca/SlimOrca-Dedup dataset.
The scores have been calculated using the best static multilingual embedding model (for my needs): sentence-transformers/static-similarity-mrl-multilingual-v1 for faster… See the full description on the dataset page:
https://huggingface.co/datasets/MarcGrumpyOlejak/slimorca_dedup_german_experimental-scored.