Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
mllm-mmr1-gt-qwen25vl7b-mmupt-full – AI Model by q1716523669 | AlphaNeural AI
You can deploy this model and start earning money today!
q1716523669
/
mllm-mmr1-gt-qwen25vl7b-mmupt-full
like
0
transformers
safetensors
grpo
gt-reward
multimodal
Qwen/Qwen2.5-VL-7B-Instruct
finetune
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Qwen2.5-VL-7B - GT-Reward on MMR1 (mmupt recipe)
GRPO with ground-truth rewards, MMR1-Math-RL-Data-v0, 481 steps (mmupt recipe: beta=0.01, 10 generations/prompt), for the big-tier GT row.
folder
step
note
best/
140
best in-loop val (eval_reward = 0.5724)
endpoint/
481
end of training
training/
holds
best_metric.json
, both
trainer_state
files,
train.log
, and the Slurm joblog.
Distinguishes from the older same-name cell by step count: this run is 481 steps (mmupt), the earlier GT is 722 steps.