Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
dpo-qwen-expG-lengthfix-adapter – AI Model by Umezaki | AlphaNeural AI
You can deploy this model and start earning money today!
Umezaki
/
dpo-qwen-expG-lengthfix-adapter
like
0
peft
safetensors
dpo
lora
qwen
structeval
text-generation
conversational
en
Qwen/Qwen3-4B-Instruct-2507
adapter
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
DPO LoRA Adapter (ExpL) for Qwen3-4B (StructEval)
SFT(ExpG) + DPO, r=8, alpha=8
SFT adapter merged from Exp G
DPO r=8, alpha=8
LR=5e-7, epoch=1, beta=0.1
Length bias fix: rejected truncated to match chosen token length
max_length=2048