This dataset contains the low-quality partition ($D_{\mathrm{low}}$) used to construct ElephantBench. The released benchmark is available separately at panzs19/ElephantBench.
$D_{\mathrm{low}}$ is derived from cx-cmu/repro-organic-data-72B using the RePro fastText quality score. It contains 47,850,862 English web documents in 600 JSONL.zstd shards (approximately 65 GiB compressed). Each record retains the source text, URL, quality score, and original… See the full description on the dataset page:
https://huggingface.co/datasets/panzs19/ElephantBench-Source.