Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.