Views
No views yet

| Stage | OLMoE 1B-7B |
|---|---|
| Base Model | allenai/OLMoE-1B-7B-0125 |
| SFT | allenai/OLMoE-1B-7B-0125-SFT |
| DPO | allenai/OLMoE-1B-7B-0125-DPO |
| Final Models (RLVR) | allenai/OLMoE-1B-7B-0125-Instruct |
| Reward Model (RM) | allenai/OLMoE-1B-7B-0125-RM |
pip install --upgrade git+https://github.com/huggingface/transformers.gitfrom transformers import AutoModelForCausalLM
olmo_model = AutoModelForCausalLM.from_pretrained("OLMoE-1B-7B-0125-Instruct")<|endoftext|><|user|>\nHow are you doing?\n<|assistant|>\nI'm just a computer program, so I don't have feelings, but I'm functioning as expected. How can I assist you today?<|endoftext|><|endoftext|><|user|>
How are you doing?
<|assistant|>
I'm just a computer program, so I don't have feelings, but I'm functioning as expected. How can I assist you today?<|endoftext|>tokenizer.apply_chat_template.You are OLMo 2, a helpful and harmless AI Assistant built by the Allen Institute for AI.| Benchmark (eval) | OLMoE-1B-7B-0125-Instruct | OLMoE-1B-7B-0924-Instruct | OLMoE-1B-7B-0125-DPO | OLMoE-1B-7B-0125-SFT | OLMoE-1B-7B-0924-SFT |
|---|---|---|---|---|---|
| Avg. | 45.62 | 38.44 | 45.05 | 41.76 | 37.05 |
| MMLU (CoT) | 55.08 | 54.57 | 54.93 | 55.26 | 54.32 |
| PopQA | 19.75 | 20.56 | 19.65 | 20.12 | 21.01 |
| TruthfulQA | 50.56 | 49.14 | 49.99 | 45.48 | 44.66 |
| BigBenchHard (CoT) | 38.61 | 36.78 | 37.37 | 37.31 | 36.55 |
| DROP | 47.87 | 34.48 | 48.38 | 48.57 | 34.71 |
| MATH (Flex) | 21.41 | 8.16 | 20.36 | 21.38 | 8.15 |
| GSM8K | 72.40 | 47.38 | 64.59 | 55.72 | 42.46 |
| HumanEval | 62.30 | 63.04 | 61.92 | 62.58 | 63.72 |
| HumanEval+ | 54.37 | 58.93 | 57.61 | 55.67 | 57.40 |
| IFEval | 66.36 | 45.29 | 65.62 | 56.56 | 41.22 |
| AlpacaEval | 17.99 | 7.54 | 19.50 | 5.83 | 6.38 |
| Safety (average) | 90.40 | 51.40 | 91.40 | 94.50 | 65.80 |
1@misc{muennighoff2024olmoeopenmixtureofexpertslanguage,
2 title={OLMoE: Open Mixture-of-Experts Language Models},
3 author={Niklas Muennighoff and Luca Soldaini and Dirk Groeneveld and Kyle Lo and Jacob Morrison and Sewon Min and Weijia Shi and Pete Walsh and Oyvind Tafjord and Nathan Lambert and Yuling Gu and Shane Arora and Akshita Bhagia and Dustin Schwenk and David Wadden and Alexander Wettig and Binyuan Hui and Tim Dettmers and Douwe Kiela and Ali Farhadi and Noah A. Smith and Pang Wei Koh and Amanpreet Singh and Hannaneh Hajishirzi},
4 year={2024},
5 eprint={2409.02060},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2409.02060},
9}
10@article{lambert2024tulu3,
11 title = {Tülu 3: Pushing Frontiers in Open Language Model Post-Training},
12 author = {
13 Nathan Lambert and
14 Jacob Morrison and
15 Valentina Pyatkin and
16 Shengyi Huang and
17 Hamish Ivison and
18 Faeze Brahman and
19 Lester James V. Miranda and
20 Alisa Liu and
21 Nouha Dziri and
22 Shane Lyu and
23 Yuling Gu and
24 Saumya Malik and
25 Victoria Graf and
26 Jena D. Hwang and
27 Jiangjiang Yang and
28 Ronan Le Bras and
29 Oyvind Tafjord and
30 Chris Wilhelm and
31 Luca Soldaini and
32 Noah A. Smith and
33 Yizhong Wang and
34 Pradeep Dasigi and
35 Hannaneh Hajishirzi
36 },
37 year = {2024},
38 email = {tulu@allenai.org}
39}