Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
exp022-dpo-ipo-merged – AI Model by tomofusa | AlphaNeural AI
You can deploy this model and start earning money today!
tomofusa
/
exp022-dpo-ipo-merged
like
0
transformers
safetensors
qwen3
text-generation
dpo
unsloth
qwen
alignment
conversational
en
u-10bei/dpo-dataset-qwen-cot
Qwen/Qwen3-4B-Instruct-2507
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
exp022-dpo-ipo-merged
SFT + DPO merged model. Full 16-bit weights, no adapter loading required.
Training Pipeline
SFT
: tomofusa/exp021b-blend-h-lora
DPO
: u-10bei/dpo-dataset-qwen-cot (1 epoch, lr=5e-07, beta=0.1)
DPO Configuration
Learning rate
: 5e-07
Beta
: 0.1
Loss type
: ipo
LoRA
: r=64, alpha=128
Max length
: 1024