This is a filtered version of Philip May's German paraphrase dataset.
The dataset has been filtered for the sake of convenience, since smaller devices do not support such large files.
All text pairs in the dataset are paraphrases, and are therefore labelled 1. As such, the dataset is well-suited for use in conjunction with the multiple negatives ranking loss.
As the original author suggests, the dataset has been filtered, mostly following the guidelines set by the author. Any row that doesn't… See the full description on the dataset page:
https://huggingface.co/datasets/danielheinz/telekom-backtrans-paraphrase-filtered.