This model is a continual pre-training of Llama-3.1-8B on a mix of mathematical datasets from SwallowMath and multilingual text datasets.
The model was trained to evaluate the performance of mathematical reasoning and problem-solving as part of the SwallowMath ablation experiments (experiment 2).
It was trained on 50 billion tokens using a mix of 4.8% SwallowMath (finemath-4+ rewritten) , 13.1% Code, and 82% multilingual text, following the setup described in the SwallowMath paper.
Training was performed using Megatron-LM.
Use
Generation
python
1# pip install -q transformers2from transformers import AutoModelForCausalLM, AutoTokenizer
34model ="tokyotech-llm/<model-name>"5device ="cuda"# for GPU usage or "cpu" for CPU usage67tokenizer = AutoTokenizer.from_pretrained(model)8model = AutoModelForCausalLM.from_pretrained(model).to(device)910inputs = tokenizer.encode("Solve the equation 2x + 3 = 7:", return_tensors="pt").to(device)11outputs = model.generate(inputs, max_length=100)12print(tokenizer.decode(outputs[0]))
Supercomputer: TSUBAME, Institute of Science Tokyo
Software
Megatron-LM (version core_r0.9.0) for training
lm-evaluation-harness for evaluation
BigCodeBench for code evaluation
Evaluation
The model was evaluated using the setup described in the SwallowMath paper, with the lm-evaluation-harness and BigCodeBench. Benchmarks include mathematical reasoning (GSM8K, MATH), code generation (HumanEval), and general tasks (OpenBookQA, TriviaQA, HellaSwag, SQuAD 2.0, XWINO, MMLU, BBH).
Results are reported for checkpoints at 10B, 20B, 30B, 40B, and 50B tokens.
Evaluation Results (SwallowMath experiment 2)
Tokens (B)
OpenBookQA
TriviaQA
HellaSwag
SQuAD2.0
XWINO
MMLU
HumanEval
GSM8K
BBH
MATH
10
0.3720
0.6643
0.5970
0.3443
0.9015
0.6343
0.3439
0.5603
0.5535
0.2480
20
0.3800
0.6580
0.5946
0.3428
0.8994
0.6293
0.3762
0.6156
0.5669
0.2860
30
0.3660
0.6618
0.5964
0.3470
0.9011
0.6298
0.3530
0.6262
0.6383
0.3040
40
0.3700
0.6610
0.5973
0.3535
0.9088
0.6358
0.3738
0.6422
0.6237
0.3100
50
0.3800
0.6637
0.5972
0.3537
0.9045
0.6337
0.3683
0.6535
0.6414
0.3160
Citation
bibtex
1@misc{fujii2025rewritingpretrainingdataboosts,
2 title={Rewriting Pre-Training Data Boosts LLM Performance in Math and Code},
3 author={Kazuki Fujii and Yukito Tajima and Sakae Mizuki and Hinari Shimada and Taihei Shiotani and Koshiro Saito and Masanari Ohi and Masaki Kawamura and Taishi Nakamura and Takumi Okamoto and Shigeki Ishida and Kakeru Hattori and Youmi Ma and Hiroya Takamura and Rio Yokota and Naoaki Okazaki},
4 year={2025},
5 eprint={2505.02881},
6 archivePrefix={arXiv},
7 primaryClass={cs.LG},
8 url={https://arxiv.org/abs/2505.02881},
9}