Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-162036 – AI Model by hirosan6595 | AlphaNeural AI
You can deploy this model and start earning money today!
hirosan6595
/
dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-162036
like
0
transformers
safetensors
qwen3
text-generation
dpo
unsloth
qwen
merged-16bit
conversational
en
u-10bei/dpo-dataset-qwen-cot
Qwen/Qwen3-4B-Instruct-2507
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
dpo-qwen-cot-merged-lr1e-7_b0.05_ep0.5_ga8-162036
This model is a
merged_16bit
model.
Initialization path
Base model:
Qwen/Qwen3-4B-Instruct-2507
Loaded and merged SFT LoRA into base:
hirosan6595/lora_structeval_t_qwen3_4b-21
Trained DPO LoRA on top and merged again
DPO Training Config (env)
lr: 1e-7
beta: 0.05
epochs: 0.5
grad_accum: 8
max_length: 2048
max_prompt_length: 1536