Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
autoscientist-mathcode-qlora – AI Model by Rishidar | AlphaNeural AI
You can deploy this model and start earning money today!
Rishidar
/
autoscientist-mathcode-qlora
like
0
safetensors
qwen2
autoscientist
adaption-labs
qlora
mathcode
dpo
en
Rishidar/autoscientist-mathcode-dataset
Qwen/Qwen2.5-0.5B-Instruct
finetune
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
AutoScientist Competition — Mathcode Model
Qwen2.5-0.5B-Instruct adapted for
mathcode
via Adaption Labs AutoScientist v5: 4-bit QLoRA SFT (r=32, alpha=64) then DPO (beta=0.1) on chosen/rejected pairs.
DPO reward accuracy: 0.8181818181818182
DPO reward margin: 8.761996030807495
Dataset:
Rishidar/autoscientist-mathcode-dataset
. Also mirrored on Kaggle: rishidard/autoscientist-mathcode-qlora.