A compiled dataset of large permissively licensed reasoning traces for experiments with midtraining / annealing before RL.
Total is ~2.5B tokens (2513520233) with OLMo 2 tokenizer.
Sources:
GeneralThought-430K (removed NC license entries), 337579 examples, 696270809 tokens
OpenThoughts-114k, 113957 examples, 766590712 tokens
OpenR1-Math-220k (all), 225129 examples, 1050658712 tokens
Sript for reformatting in repo.