Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen2.5-3b-dpo-vietnamese – AI Model by Haanh1702 | AlphaNeural AI
You can deploy this model and start earning money today!
Haanh1702
/
qwen2.5-3b-dpo-vietnamese
like
0
safetensors
alignment
dpo
trl
unsloth
vietnamese
text-generation
vi
en
unsloth/Qwen2.5-3B-bnb-4bit
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5-3B-DPO-Vietnamese
Mo hinh
Qwen2.5-3B
duoc can chinh so thich (Alignment) bang thuat toan
Direct Preference Optimization (DPO)
trong khuon kho bai Lab 22 (Track 3 - AICB Program).
Thong tin huan luyen
Base Model:
unsloth/Qwen2.5-3B-bnb-4bit
SFT Dataset:
kai-foundation-models/vi-alpaca (1,000 samples)
Preference Dataset:
rgilla/ultrafeedback-binarized-preferences-cleaned (1,000 pairs)
Hyperparameters:
eta: 0.1
learning_rate: 5e-07
epochs: 1
optimizer: adamw_8bit
loss_type: sigmoid
Ket qua thuc nghiem
Final DPO Loss:
0.8484
End Chosen Reward:
-0.4988
End Rejected Reward:
-0.5331
Reward Gap:
+0.0343
Qualitative Win-rate (8 test prompts):
62.5% Thang (5/8), 37.5% Hoa (3/8), 0% Thua.