The shared ChatML instruct dataset (SmolTalk core + Dolly-15k) and 32k tokenizer used to SFT both arms of the A/B below. All-masked windows are dropped at pack time.
Two ~100M-parameter from-scratch English language models, identical architecture
and identical instruct SFT, differing only in pretraining token budget:
model
pretrain tokens
instruct val loss… See the full description on the dataset page:
https://huggingface.co/datasets/Aeryx-ai/lm-100m-en-sft-data.