A Llama 3.2-tokenized re-windowing of JackHsieh/statML-arxiv-40M-20M
(originally tokenized with Qwen3), the Llama analogue of
JackHsieh/statML-arxiv-40M-20M-olmo3.
Each row is a contiguous span of exactly 4096 tokens under the Llama 3.2 tokenizer
(meta-llama/Llama-3.2-3B, byte-identical to the 1B tokenizer, vocab 128 256), anchored at the same
character offset as the corresponding Qwen3 window of the same paper.
train: 9,728 sequences (39.8M tokens)… See the full description on the dataset page:
https://huggingface.co/datasets/JackHsieh/statML-arxiv-40M-20M-llama32.