Views
No views yet

| Method | MMLU | GSM8k | BBH | TydiQA | Codex | Squad | AlpacaEval | Average |
|---|---|---|---|---|---|---|---|---|
| Random (unbal.) | 61.6 | 81.2 | 66.8 | 71.1 | 76.4 | 89.7 | 75.6 | 74.6 |
| Random (bal.) | 62.1 | 76.0 | 68.6 | 68.8 | 87.2 | 87.4 | 72.4 | 74.7 |
| Tulu 3 SFT | 62.2 | 74.3 | 68.2 | 67.4 | 83.8 | 85.5 | 71.9 | 73.3 |
| RDS+ (this model) | 62.5 | 77.6 | 66.6 | 72.1 | 83.8 | 90.2 | 80.2 | 76.1 |
| RDS+ - Arena Hard | 57.0 | 78.7 | 59.7 | 49.4 | 75.7 | 66.3 | 84.5 | 67.3 |
<|user|>
Your message here!
<|assistant|><|assistant|>, this can affect generation quality quite a bit.
We have included a chat template in the tokenizer implementing this template.@misc{ivison2025data,
title={{Practical Large-Scale Data Selection for Instruction Tuning}},
author={{Hamish Ivison and Muru Zhang and Faeze Brahman and Pang Wei Koh and Pradeep Dasigi}}
year={2025},
url={https://arxiv.org/abs/2503.01807},
eprint={2503.01807},
archivePrefix={arXiv},
primaryClass={cs.CL}
}