This dataset is a curated, tokenized pretraining mixture designed specifically for training MiniModel-series small language models. It was tokenized using the Mistral-7B-Instruct-v0.3 tokenizer (vocab size: 32,768), which is included in the MiniModel-200M-Base repository.
For training code, data loading utilities, and full reproducibility (including the training script), see the official GitHub repository:🔗… See the full description on the dataset page:
https://huggingface.co/datasets/xTimeCrystal/TinyCorpus-v2.