Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
slm-125m-dpo – AI Model by thesreedath | AlphaNeural AI
You can deploy this model and start earning money today!
thesreedath
/
slm-125m-dpo
like
0
safetensors
llama
legal
dpo
preference-optimization
slm
en
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
slm-125m-dpo
125M legal SLM, QA-SFT then
DPO
(beta=0.1) on an AI-feedback preference set (chosen vs. deliberately-flawed rejected). Closed-book. Reference = the QA-SFT model.