Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
qwen2.5-3b-dpo-grpo – AI Model by Phantomcloak19 | AlphaNeural AI
You can deploy this model and start earning money today!
Phantomcloak19
/
qwen2.5-3b-dpo-grpo
like
0
transformers
safetensors
qwen2
text-generation
horus-llm
dpo-grpo
qwen2.5-3b
conversational
Qwen/Qwen2.5-3B-Instruct
finetune
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
qwen2.5-3b — DPO-GRPO
Merged full-precision model after the
DPO-GRPO
phase of the LLMPR sequential training pipeline (SFT → DPO → Safety-GRPO).
Field
Value
Base model
Qwen/Qwen2.5-3B-Instruct
Phase
DPO-GRPO
Short name
qwen2.5-3b
Generated
2026-06-30 03:11 UTC