Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
slm-125m-ppo – AI Model by rahulreddyhanu | AlphaNeural AI
You can deploy this model and start earning money today!
rahulreddyhanu
/
slm-125m-ppo
like
0
safetensors
llama
legal
finance
rlaif
ppo
text-generation
en
rahulreddyhanu/slm-125m-legal-financial-sft
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
slm-125m-ppo
PPO/RLAIF of the 125M legal/financial SFT model against a reward model (125M SFT + scalar head, Bradley-Terry on the preference triplets). 60-step PPO loop; mean reward -1.27 -> -0.64.
Base:
rahulreddyhanu/slm-125m-legal-financial-sft
. Part of an RLAIF demo (DPO + PPO) on small legal/financial models.