Subset of JackHsieh/statML-arxiv. Each row is one randomly
sampled contiguous window of exactly 1_024 Llama 3.2 tokens (meta-llama/Llama-3.2-3B) from a
distinct paper: papers with at least 1_024 tokens (Llama3.2_token_count) are taken in
yymm-descending order, and each window is length-nudged by up to 16 tokens (start held fixed), resampling the position only if that fails, until its decoded text re-encodes to
exactly 1_024 tokens. start_index is the window's offset in the source paper's token… See the full description on the dataset page:
https://huggingface.co/datasets/JackHsieh/statML-arxiv-42M-11M-L-1024-llama32.