Dataset Card for yaakov/wikipedia-de-splits
Dataset Description
The only goal of this dataset is to have random German Wikipedia articles at
various dataset sizes: Small datasets for fast development and large datasets for statistically relevant measurements.
For this purpose, I loaded the 2665357 articles in the test set of the pre-processed German Wikipedia dump from 2022-03-01, randomly permuted the articles and created splits of sizes 2**n: 1, 2, 4, 8, .... The… See the full description on the dataset page: https://huggingface.co/datasets/yaakov/wikipedia-de-splits.