Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
exp018-dpo-merged – AI Model by tomofusa | AlphaNeural AI | AlphaNeural AI
You can deploy this model and start earning money today!
tomofusa
/
exp018-dpo-merged
like
0
transformers
safetensors
qwen3
text-generation
dpo
unsloth
qwen
alignment
conversational
en
u-10bei/dpo-dataset-qwen-cot
Qwen/Qwen3-4B-Instruct-2507
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
exp018-dpo-merged
SFT + DPO merged model. Full 16-bit weights, no adapter loading required.
Training Pipeline
SFT
: tomofusa/exp015-blend-h-lora
DPO
: u-10bei/dpo-dataset-qwen-cot (1 epoch, lr=5e-07, beta=0.01)
DPO Configuration
Learning rate
: 5e-07
Beta
: 0.01
Loss type
: ipo
LoRA
: r=64, alpha=128
Max length
: 1024