Private research corpus with 32,000,010,072 globally
exact-deduplicated train tokens plus 328,933,246 held-out
tokens. Data are stored as EOS-delimited little-endian uint16 binaries with
aligned Parquet provenance.
This repository combines ODC-By FinePDFs-Edu, CC-BY-4.0 DCLM, ODC-By
SmolLM/FineMath sources, Apache-2.0 UltraX subject to its upstream terms,
per-file permissively licensed Stack-Edu code subject to The Stack v2 terms,
StarCoder2… See the full description on the dataset page:
https://huggingface.co/datasets/glouriousgautam/lilm1-pretrain-mix-32b.