German4All is a synthetic data corpus consisting of 3 datasets. Each dataset consists of German Wikipedia paragraphs that are paraphrased in five different complexity levels. The 3 datasets are:
German4All-Main (subfolder "main"): The main synthetic dataset containing 25,459 elements, each featuring an
original text along with its five-level paraphrases.
German4All-Main-old (subfolder… See the full description on the dataset page:
https://huggingface.co/datasets/tum-nlp/German4All-Corpus.