MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code… See the full description on the dataset page:
https://huggingface.co/datasets/NP235/MathCode-Pile.