Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen2.5-1.5b-dpo-truthful – AI Model by sak3E | AlphaNeural AI
You can deploy this model and start earning money today!
sak3E
/
qwen2.5-1.5b-dpo-truthful
like
0
safetensors
dpo
qlora
truthfulness
alignment
en
Qwen/Qwen2.5-1.5B-Instruct
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5-1.5B — DPO Fine-tuned for Truthfulness
Fine-tuned from
Qwen/Qwen2.5-1.5B-Instruct
using
Direct Preference Optimization (DPO)
on
jondurbin/truthy-dpo-v0.1
to reduce hallucinations and improve factual accuracy.
Training details
Method: DPO (beta=0.2) + QLoRA (4-bit NF4, rank 16)
Optimizer: adamw_torch
Best config: 500 samples, 1 epoch, LR: 2e-5, beta: 0.2
Evaluated with AlpacaEval (helpful_base, 15 samples) — win rate: 76.7%