Views
No views yet
[!IMPORTANT] 🔥 Excited to find ourDART-Math-DSMath-7B(Prop2Diff) comparable to the AIMO winner NuminaMath-7B on CoT, but based solely on MATH & GSM8K prompt set, leaving much room to improve! Besides, ourDARTmethod is also fully compatible with tool-integrated reasoning. Find more details and join the discussion under this X thread!
DART-MathDART-Math models achieve performance superior or competitive to previous SOTAs on 2 in-domain and 4 challenging out-of-domain mathematical reasoning benchmarks, despite using much smaller datasets and no proprietary model like GPT-4.| Model | MATH | GSM8K | College | DM | Olympiad | Theorem | AVG |
|---|---|---|---|---|---|---|---|
| GPT-4 (0314) | 52.6 | 94.7 | 24.4 | -- | -- | -- | -- |
| Llama-3-70B-MetaMath | 44.9 | 88.0 | 31.9 | 53.2 | 11.6 | 21.9 | 41.9 |
DART-Math-Llama-3-70B (Uniform) | 54.9 | 90.4 | 38.5 | 64.1 | 19.1 | 27.4 | 49.1 |
DART-Math-Llama-3-70B (Prop2Diff) | 56.1 | 89.6 | 37.9 | 64.1 | 20.0 | 28.2 | 49.3 |
| DeepSeekMath-7B-MetaMath | 43.7 | 81.8 | 33.7 | 53.0 | 13.6 | 23.2 | 41.5 |
| DeepSeekMath-7B-RL | 53.1 | 88.4 | 41.3 | 58.3 | 18.7 | 35.9 | 49.3 |
DART-Math-DSMath-7B (Uniform) | 52.9 | 88.2 | 40.1 | 60.2 | 21.3 | 32.5 | 49.2 |
DART-Math-DSMath-7B (Prop2Diff) | 53.6 | 86.8 | 40.7 | 61.6 | 21.7 | 32.2 | 49.4 |
| Mistral-7B-MetaMath | 29.8 | 76.5 | 19.3 | 28.0 | 5.9 | 14.0 | 28.9 |
DART-Math-Mistral-7B (Uniform) | 43.5 | 82.6 | 26.9 | 42.0 | 13.2 | 16.4 | 27.4 |
DART-Math-Mistral-7B (Prop2Diff) | 45.5 | 81.1 | 29.4 | 45.1 | 14.7 | 17.0 | 38.8 |
| Llama-3-8B-MetaMath | 32.5 | 77.3 | 20.6 | 35.0 | 5.5 | 13.8 | 30.8 |
DART-Math-Llama-3-8B (Uniform) | 45.3 | 82.5 | 27.1 | 48.2 | 13.6 | 15.4 | 38.7 |
DART-Math-Llama-3-8B (Prop2Diff) | 46.6 | 81.1 | 28.8 | 48.0 | 14.5 | 19.4 | 39.7 |
DART-Math GitHub repository.DART-Math models use the Alpaca prompt template:
Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n###Instruction:\n{query}\n\n### Response:\n
DARS) to the MATH and GSM8K training sets.DARS tackle severe biases towards easy queries, with frequent failures to generate any correct response for the most challenging queries, in previous datasets.DART-Math-Hard / DART-Math-Uniform for more details.DART-Math-Hard & DART-Math-Uniform,
leading to DART-Math (Prop2Diff) & DART-Math (Uniform) respectively.| Base Model | Max. L.R. | # of Epochs | # of Grad. Acc. Steps | # of A100 GPUs |
|---|---|---|---|---|
| Mistral-7B | 1e-5 | 3 | 1 | 8 |
| Llama3-8B | 5e-5 | 1 | 2 | 8 |
| Llama3-70B | 2e-5 | 1 | 1 | 32 |
| DeepSeekMath-7B | 5e-5 | 3 | 1 | 8 |
1e-6,5e-6,1e-5,2e-5,5e-5,1e-4 according to the MATH performance after training on MMIQC for 1 epoch, except for Llama3-70B that is so expensive to search for that we derive from Llama3-8B’s learning rate in analogy to the relationship of (per-training) learning rates between Llama2-7B and Llama2-70B (~2:1).sliding_window by default following the newest Mistral-7B-Instruct (Flash Attention 2 does not support sliding_window and XFormer backend in vLLM has throughput ~10% lower in our experiments.)1@article{tong2024dartmath,
2 title={DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving},
3 author={Yuxuan Tong and Xiwen Zhang and Rui Wang and Ruidong Wu and Junxian He},
4 year={2024},
5 eprint={2407.13690},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2407.13690},
9}