A test dataset of one million identical rows, duplicated from a single source row. Useful for stress-testing data loaders, batch I/O, and Hub upload pipelines.
The data is sharded into batch files of 100,000 rows each, named batch_000.csv through batch_009.csv. Each file contains the… See the full description on the dataset page:
https://huggingface.co/datasets/burtenshaw/1-million-rows.