A 10M-row subset of the stage-1 pretraining mixture from DreamCoder, sampled according to their original mixture weights:
Each row has a text column (the example) and a source column (which mixture component it came from). Rows are globally… See the full description on the dataset page:
https://huggingface.co/datasets/EER6/adlmc-stage1-10M.