This is the original and unchanged german translated dataset (train split only) in original order from DiscoResearch/germanrag with added cosine-similarity scores.
The scores between 'question' and 'answer' have been calculated using the best static multilingual embedding model (for my needs): sentence-transformers/static-similarity-mrl-multilingual-v1 for faster distinction if an answer corresponds to a query upon the content.
If you want to filter negative answers… See the full description on the dataset page:
https://huggingface.co/datasets/MarcGrumpyOlejak/germanrag-scored.