This is a recipe, not a corpus. It contains pinned source revisions and deterministic
scripts that reconstruct a ~400M-token English pretraining blend byte-for-byte, plus the
provenance record from the build that produced it. No corpus text, tokenizer, or model
weights are stored in this repository.
The blend is built for episod/tt-tnt (public), a
hand-rolled, nanollama3-lineage, Llama-3-architecture language model trained with tt-metal's
ttml trainer on… See the full description on the dataset page:
https://huggingface.co/datasets/episod/tt-tnt-corpus.