Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen2.5-1.5b-dpo-truthful-mac – AI Model by sak3E | AlphaNeural AI
You can deploy this model and start earning money today!
sak3E
/
qwen2.5-1.5b-dpo-truthful-mac
like
0
safetensors
dpo
lora
truthfulness
mac-m4
en
Qwen/Qwen2.5-1.5B-Instruct
adapter
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5-1.5B — DPO Fine-tuned for Truthfulness (Mac M4)
Fine-tuned from
Qwen/Qwen2.5-1.5B-Instruct
using
Direct Preference Optimization (DPO)
on
jondurbin/truthy-dpo-v0.1
to reduce hallucinations.
Training details
Method: DPO (beta=0.1) + LoRA (rank 16)
Hardware: Apple Silicon Mac M4 (MPS)
No quantization (float16)
Samples: 50, Epochs: 1