Pretokenized training mixture for compact TR-HASH language-model research.
The Dataset Viewer displays one summary row per source. The actual training data
is stored as packed uint16 token shards under corpora//tokens-*.bin.
Each sequence contains 1,024 token IDs produced by the project 32K tokenizer.