Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
q0-gsm8k-1_5b-mopd – AI Model by Pradheep1647 | AlphaNeural AI
You can deploy this model and start earning money today!
Pradheep1647
/
q0-gsm8k-1_5b-mopd
like
0
transformers
safetensors
qwen2
text-generation
gsm8k
distillation
mopd
q0
conversational
Qwen/Qwen2.5-1.5B-Instruct
finetune
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
q0-gsm8k-1.5b-mopd
Qwen2.5-1.5B-Instruct fine-tuned with
MOPD
(multi-teacher on-policy distillation) on GSM8K, distilled from a q0 cyclic-trajectory GRPO mixture.
Base:
Qwen/Qwen2.5-1.5B-Instruct
(revision
989aa7980e4cf806f80c7fef2b1adb7bc71aa306
)
Method:
MOPD, teacher_count=2 (top-2 q0 snapshots: cycle02 + cycle03, uniformly averaged), student = 4-bit QLoRA (NF4, r=8, alpha=16).
Selection:
step 192 by GSM8K validation pass@8 (256-example slice, seed 42).
GSM8K validation:
pass@1 49.07%, pass@4 76.29%, pass@8 84.77% (256-example deterministic slice).
MATH-500 test:
pass@1 16.40%, pass@4 35.00% (500 problems, four samples each).
Format:
LoRA adapter merged into 16-bit base weights (deployable directly).
Trained on an RTX 4060 Laptop (8 GB) with 4-bit QLoRA. This is the 1.5B checkpoint (the repository's 1B-scale target), not a literal 1.0B model.