Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
tiny-think-dpo-math-stem-apo_zero-beta1-lr3e-6-e1-bs8 – AI Model by Shekswess | AlphaNeural AI
You can deploy this model and start earning money today!
Shekswess
/
tiny-think-dpo-math-stem-apo_zero-beta1-lr3e-6-e1-bs8
like
0
transformers
safetensors
llama4_text
text-generation
generated_from_trainer
dpo
trl
conversational
2305.18290
Shekswess/tiny-think-sft-math-stem-loss-nll-bf16-lr2e-5-e2-bs8
finetune
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Model Card for tiny-think-dpo-math-stem-apo_zero-beta1-lr3e-6-e1-bs8
This model is a fine-tuned version of
Shekswess/tiny-think-sft-math-stem-loss-nll-bf16-lr2e-5-e2-bs8
. It has been trained using
TRL
.
Training procedure
This model was trained with DPO, a method introduced in
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
.
Framework versions
TRL: 0.26.2
Transformers: 4.57.5
Pytorch: 2.9.0+cu128
Datasets: 4.5.0
Tokenizers: 0.22.2