Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
slm-125m-dpo – AI Model by rahulreddyhanu | AlphaNeural AI
You can deploy this model and start earning money today!
rahulreddyhanu
/
slm-125m-dpo
like
0
safetensors
llama
legal
finance
rlaif
dpo
text-generation
en
rahulreddyhanu/slm-125m-legal-financial-sft
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
slm-125m-dpo
DPO alignment of the 125M legal/financial SFT model on 498 AI-feedback preference triplets (GPT-4o-mini chosen vs on-policy rejected, Gemini 2.5 Flash judge). beta=0.1, 1 epoch.
Base:
rahulreddyhanu/slm-125m-legal-financial-sft
. Part of an RLAIF demo (DPO + PPO) on small legal/financial models.