Views
No views yet
| Model | LC. Win Rate | Win Rate | Avg. Length |
|---|---|---|---|
| Gemma-2-9B-SPPO Iter1 | 48.70 | 40.76 | 1669 |
| Gemma-2-9B-SPPO Iter2 | 50.93 | 44.64 | 1759 |
| Gemma-2-9B-SPPO Iter3 | 53.27 | 47.74 | 1803 |
@misc{wu2024self,
title={Self-Play Preference Optimization for Language Model Alignment},
author={Wu, Yue and Sun, Zhiqing and Yuan, Huizhuo and Ji, Kaixuan and Yang, Yiming and Gu, Quanquan},
year={2024},
eprint={2405.00675},
archivePrefix={arXiv},
primaryClass={cs.LG}
}