This dataset contains the final merged 20B-token scratch-pretraining mixture used for OpenAster-1. Files are released as line-preserving jsonl.zst shards. Each record contains text, source, and language.
DCLM Edu: 4B target tokens, English educational web.
OpenCSG FineWeb Edu 4/5: 5B target tokens, Chinese educational web.
FineMath 4+: 2B target tokens, English math corpus.
FineWeb Edu EN: 4B target tokens, English… See the full description on the dataset page:
https://huggingface.co/datasets/binichallein/OpenAster-1-data.