Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
slm-gemma-2b-dpo – AI Model by rahulreddyhanu | AlphaNeural AI
You can deploy this model and start earning money today!
rahulreddyhanu
/
slm-gemma-2b-dpo
like
0
safetensors
gemma2
legal
finance
rlaif
dpo
text-generation
conversational
en
thesreedath/slm-gemma-2b-qa
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
slm-gemma-2b-dpo
LoRA DPO alignment of the Gemma-2B legal QA SFT model on 470 AI-feedback preference triplets (adapter merged). beta=0.1.
Base:
thesreedath/slm-gemma-2b-qa
. Part of an RLAIF demo (DPO + PPO) on small legal/financial models.