Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
Phi-3-mini-4k-instruct-bandit-rm-700k – AI Model by yyqoni | AlphaNeural AI
You can deploy this model and start earning money today!
yyqoni
/
Phi-3-mini-4k-instruct-bandit-rm-700k
like
0
transformers
safetensors
phi3
text-classification
custom_code
hendrydong/preference_700K
2501.02790
microsoft/Phi-3-mini-4k-instruct
finetune
mit
autotrain_compatible
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
This is the bandit reward model introduced in the preprint
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Models
(
https://arxiv.org/abs/2501.02790
). For more details, please visit our repository at
https://github.com/yinyueqin/DenseRewardRLHF-PPO
.