Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
mllm-cogrpo-heter-qwen25vl-7b-x-internvl35-8b-mmr1-mmupt-groupA-qwen25vl-7b – AI Model by q1716523669 | AlphaNeural AI
You can deploy this model and start earning money today!
q1716523669
/
mllm-cogrpo-heter-qwen25vl-7b-x-internvl35-8b-mmr1-mmupt-groupA-qwen25vl-7b
like
0
safetensors
qwen2_5_vl
co-rl
co-learn
cogrpo
mllm
mmr1
reasoning
image-text-to-text
conversational
Qwen/Qwen2.5-VL-7B-Instruct
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
q1716523669/mllm-cogrpo-heter-qwen25vl-7b-x-internvl35-8b-mmr1-mmupt-groupA-qwen25vl-7b
Cross-family
co-learning
(co-GRPO-DP): trained jointly with peer
OpenGVLab/InternVL3_5-8B-HF
via majority-vote cross-labeling — no ground-truth labels are used.
Base: Qwen/Qwen2.5-VL-7B-Instruct · Peer: OpenGVLab/InternVL3_5-8B-HF
Method: co-GRPO-DP (heterogeneous co-learning, rendezvous cross-pseudo-labels)
Dataset: mmr1 (~8k), 1 epoch = 722 steps
Checkpoint:
best
(best_model)
4-bench avg: **48.28 ** (MathVision / MathVerse / MathVista / We-Math; prompt=answer, greedy T=0, mathruler)
WandB project: co-rl-mllm-mmr1