Mar.31 2025 Added instruct fine-tuning and reasoning dataset in the same ELI5 style. Take a look!
Mar.31 2025 See my new model.
Inspired by the paper TinyHelen's First Curriculum, we present a Chinese version of the LLM-simplified training corpus. This dataset is converted from high-quality Chinese and English web crawls for training baby-size (<100M) language models.
Adult-talking
北京市财政局、北京海关、国家税务总局北京市税务局、北京市国际服务贸易事务中心:… See the full description on the dataset page:
https://huggingface.co/datasets/fzmnm/TinyHelen-zh.