This datasets includes all corpora that were used for pretraining the German DBMDZ BERT Models.
It consists of Wikipedia dump and corpora from OPUS:
Filename
Description
Creation Date
File Size
dewiki.txt
Wikipedia Dump
May 2019
5.1GB
eubookshop.txt
OPUS EUbookshop
November 2018
2.2GB
news.2018.txt
OPUS News corpora
January 2019
4.1GB
opensubtitles.txt
OPUS OpenSubtitles
November 2018
1.3GB
paracrawl.txt
OPUS ParaCrawl
November 2018
3.1GB… See the full description on the dataset page:
https://huggingface.co/datasets/stefan-it/german-dbmdz-bert-corpus.