Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
slm-gemma-2b-dpo2 – AI Model by thesreedath | AlphaNeural AI
You can deploy this model and start earning money today!
thesreedath
/
slm-gemma-2b-dpo2
like
0
safetensors
gemma2
legal
dpo
gemma-2
preference-optimization
en
gemma
4-bit
bitsandbytes
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
slm-gemma-2b-dpo2
Gemma 2 2B: gemma-2-2b-it -> QA SFT -> instruction SFT ->
DPO v2
(QLoRA, beta=0.1) on ~4k AI-feedback preference pairs spanning closed-book QA AND instruction-following failure modes, prompts held out of both SFT sets.