Tokenized pretraining corpus for Merlin, a 3B LLM for agentic coding on Apple Silicon.
Note: v0 does not yet include agentic traces. Tokenizer and corpus will be retrained after trace generation.
corpus_train.bin + corpus_val.bin, uint16, shape [N, 6144]
90/10 train/val split at document boundaries (seed=42 shuffle)
~1.19B tokens total, 100% packing efficiency
Fits entirely in H100 HBM (~2.4GB)
Tokenizer: tsuberim/merlin-tokenizer-v0… See the full description on the dataset page:
https://huggingface.co/datasets/tsuberim/merlin-corpus-v0.