Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwama-0.5b-skywork-pref-dpo-trl-v2 – AI Model by lblaoke | AlphaNeural AI
You can deploy this model and start earning money today!
lblaoke
/
qwama-0.5b-skywork-pref-dpo-trl-v2
like
0
safetensors
qwen2
Skywork/Skywork-Reward-Preference-80K-v0.1
turboderp/Qwama-0.5B-Instruct
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
learning_rate: 5.0e-6
num_train_epochs: 4
per_device_train_batch_size: 2
gradient_accumulation_steps: 8