document-corpus-v3-open is the redistribution-compatible slice of the exact
byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT
experiments. It contains 869,739 filtered documents and
2.192 GB of UTF-8 text before Parquet compression.
This is not the complete internal document-corpus-v3. Restricted,
unknown-license, and share-alike sources were excluded conservatively. Every
included row comes from an upstream dataset whose card… See the full description on the dataset page:
https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.