Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
R1-VL-7B – AI Model by jingyiZ00 | AlphaNeural AI
You can deploy this model and start earning money today!
jingyiZ00
/
R1-VL-7B
like
0
transformers
safetensors
qwen2_vl
image-to-text
image-text-to-text
conversational
HuanjinYao/Mulberry-SFT
2503.12937
Qwen/Qwen2-VL-7B-Instruct
finetune
apache-2.0
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
R1-VL-7B
R1-VL-7B is a reasoning model trained with step-wise group relative policy optimization (StepGRPO).
Paper:
https://arxiv.org/pdf/2503.12937
Github:
https://github.com/jingyi0000/R1-VL
Base model:
https://huggingface.co/Qwen/Qwen2-VL-7B-Instruct