This is the original and unchanged german translated dataset (train split only) in original order from aari1995/ultradistil-intel-orca-dpo-de with added cosine-similarity scores.
Only for 'input' and 'chosen' scores have been calculated. The scores have been calculated using the best static multilingual embedding model (for my needs): sentence-transformers/static-similarity-mrl-multilingual-v1 for faster distinction if an answer corresponds to a query upon the content.… See the full description on the dataset page:
https://huggingface.co/datasets/MarcGrumpyOlejak/ultradistil-intel-orca-dpo-de-scored.