This is a filtered and deduplicated version of the german subset of the 23.01 OSCAR Corpus, a large, crawled, and processed text dataset
curated by the OSCAR project (Open Super-large Crawled Aggregated coRpus).
OSCAR 23.01 is the January 2023 version of the OSCAR Corpus based on the November/December 2022 dump of Common Crawl.
While being quite similar to OSCAR 22.01, it contains several new features, including KenLM-based adult content detection… See the full description on the dataset page:
https://huggingface.co/datasets/bjoernp/oscar2023_deduped_filtered_1.1.