LittleCurriculum is an ~88B-token English pretraining corpus derived from FineWeb-Edu. It is filtered to align with U.S. Common Core standards for grades K–5, removing documents containing academic concepts and skills characteristic of later grades.
It is the training corpus for the LittleLearner models, designed to study language models under a controlled knowledge boundary. See the LittleLearner paper for methodology, validation, and experiments.… See the full description on the dataset page:
https://huggingface.co/datasets/littlelearner/LittleCurriculum.