Frozen, per-model binned train/test splits for the MRCR experiments in Randomized YaRN
Improves Length Generalization for Long-Context Reasoning (Mehta, Yin, Durrett). MRCR samples
are sorted into length bins by token count, which is tokenizer-dependent — so each model gets
its own config. Published as fixed bytes so the paper reproduces independently of
transformers versions.
config
tokenizer
train
test