Views
No views yet
| Model | LC. Win Rate | Win Rate |
|---|---|---|
| Llama-3-Base-8B-SFT-DPO | 18.20 | 15.50 |
| Llama-3-Base-8B-DICE-Iter1 | 25.08 | 25.77 |
| Llama-3-Base-8B-DICE-Iter2 | 27.55 | 30.99 |
1@article{chen2024bootstrapping,
2 title={Bootstrapping Language Models with DPO Implicit Rewards},
3 author={Chen, Changyu and Liu, Zichen and Du, Chao and Pang, Tianyu and Liu, Qian and Sinha, Arunesh and Varakantham, Pradeep and Lin, Min},
4 journal={arXiv preprint arXiv:2406.09760},
5 year={2024}
6}