Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
dpo-qwen-cot-v4-merged – AI Model by y-tani | AlphaNeural AI
You can deploy this model and start earning money today!
y-tani
/
dpo-qwen-cot-v4-merged
like
0
transformers
safetensors
qwen3
text-generation
dpo
qwen
alignment
conversational
en
u-10bei/dpo-dataset-qwen-cot
Qwen/Qwen3-4B-Instruct-2507
quantized
apache-2.0
text-generation-inference
endpoints_compatible
4-bit
bitsandbytes
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
y-tani/dpo-qwen-cot-v4-merged
This model is a fine-tuned version of
Qwen/Qwen3-4B-Instruct-2507
using a two-stage pipeline:
SFT (Supervised Fine-Tuning)
: LoRA adapter from
y-tani/lora_structeval_t_qwen3_4b_v4
DPO (Direct Preference Optimization)
: Applied on top of the SFT model
This repository contains the
full-merged 16-bit weights
. No adapter loading is required.
Training Configuration
Base model
: Qwen/Qwen3-4B-Instruct-2507
SFT adapter
: y-tani/lora_structeval_t_qwen3_4b_v4
DPO dataset
: u-10bei/dpo-dataset-qwen-cot
Method
: SFT → DPO (two-stage)
Epochs
: 1
Learning rate
: 1e-7
Beta
: 0.1
Max sequence length
: 1024
DPO LoRA Config
: r=8, alpha=16 (merged into SFT base)
Sources & License (IMPORTANT)
Training Data
: [u-10bei/dpo-dataset-qwen-cot]
License
: MIT License. (As per dataset terms).
Compliance
: Users must follow the original base model's license terms.