Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
exp43-dpo-on-exp40 – AI Model by ogwata | AlphaNeural AI
You can deploy this model and start earning money today!
ogwata
/
exp43-dpo-on-exp40
like
0
transformers
safetensors
qwen3
text-generation
dpo
unsloth
qwen
alignment
conversational
en
u-10bei/dpo-dataset-qwen-cot
ogwata/exp40-lr1e5-wd01-merged
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
exp43-dpo-on-exp40
This model is a fine-tuned version of
ogwata/exp40-lr1e5-wd01-merged
using
Direct Preference Optimization (DPO)
via the
Unsloth
library.
This repository contains the
full-merged 16-bit weights
. No adapter loading is required.
Training Configuration
Base model
: ogwata/exp40-lr1e5-wd01-merged
Method
: DPO (Direct Preference Optimization)
Epochs
: 1
Learning rate
: 7e-07
Beta
: 0.2
Max sequence length
: 1024
LoRA Config
: r=8, alpha=16 (merged into base)