Reproducible encoded training artifact shared by the MOSAIC D384 model
family. It contains the fixed 10M-word Strict-Small corpus representation used
for a controlled 100M-word exposure schedule.
packed token and word arrays for training and validation;
document offsets and token frequencies;
document-complexity metadata;
lexical recombination neighbors used by the variation-set arms;
manifest.json with source and artifact… See the full description on the dataset page:
https://huggingface.co/datasets/miguelcsx/babylm-2026-mosaic-bpe-corpus.