Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
zephyr-7b-gemma-rpo-avg – AI Model by ZHLiu627 | AlphaNeural AI
You can deploy this model and start earning money today!
ZHLiu627
/
zephyr-7b-gemma-rpo-avg
like
0
safetensors
gemma
argilla/dpo-mix-7k
2405.16436
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
This models uses
OpenRLHF Codebase
for the average loss with the method
Regularized-Preference-Optimization
. The SFT loss coefficient is
0.2
. The relevant paper is (
https://arxiv.org/abs/2405.16436
).