Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen3-4b-structured-sft-dpo – AI Model by ryusuke009 | AlphaNeural AI
You can deploy this model and start earning money today!
ryusuke009
/
qwen3-4b-structured-sft-dpo
like
0
transformers
safetensors
qwen3
text-generation
sft
dpo
unsloth
qwen
conversational
en
u-10bei/structured_data_with_cot_dataset_512_v2
u-10bei/dpo-dataset-qwen-cot
Qwen/Qwen3-4B-Instruct-2507
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen3-4b-structured-sft-dpo
Two-stage fine-tuned model: SFT + DPO
Training
Stage 1 (SFT): QLoRA on structured_data_with_cot_dataset_512_v2 (LR=2e-6, Epochs=2, LoRA r=64)
Stage 2 (DPO): DPO on dpo-dataset-qwen-cot (LR=1e-07, Epochs=1, Beta=0.1, LoRA r=8)