Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
rlhflow-llama-3-sft-8b-v2-bandit-ppo-60k – AI Model by yyqoni | AlphaNeural AI
You can deploy this model and start earning money today!
yyqoni
/
rlhflow-llama-3-sft-8b-v2-bandit-ppo-60k
like
0
transformers
safetensors
llama
text-generation
conversational
argilla/ultrafeedback-binarized-preferences-cleaned
2501.02790
RLHFlow/LLaMA3-SFT-v2
finetune
mit
autotrain_compatible
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
This is the bandit reward based ppo model introduced in the preprint
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Models
(
https://arxiv.org/abs/2501.02790
). For more details, please visit our repository at
https://github.com/yinyueqin/DenseRewardRLHF-PPO
.