Views
No views yet

llama_3_license.txt.| Model | Score | Chat | Chat Hard | Safety | Reasoning |
|---|---|---|---|---|---|
| Llama 3.1 Tulu 2 8b UF RM (this model) | 73.3 | 98.0 | 59.6 | 60.6 | 74.7 |
| Llama 3.1 Tulu 2 70b UF RM | 70.2 | 96.4 | 56.4 | 65.8 | 62.3 |
ultrafeedback_mean_aspects split.<|user|>
Your message here!
<|assistant|><|assistant|>, this can affect generation quality quite a bit.
We have included a chat template in the tokenizer implementing this template.@misc{ivison2024unpacking,
title={{Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback}},
author={{Hamish Ivison and Yizhong Wang and Jiacheng Liu and Ellen Wu and Valentina Pyatkin and Nathan Lambert and Yejin Choi and Noah A. Smith and Hannaneh Hajishirzi}}
year={2024},
eprint={2406.09279},
archivePrefix={arXiv},
primaryClass={cs.CL}
}