Beta
Explore
Marketplace
Neural Labs
Playground
Wallet
Docs
simple_GRPO – AI Model by whysue | AlphaNeural AI
You can deploy this model and start earning money today!
whysue
/
simple_GRPO
like
0
safetensors
qwen2
text-generation-inference
question-answering
en
openai/gsm8k
Qwen/Qwen2.5-7B-Instruct
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
license: mit base_model:
Qwen/Qwen2.5-7B
参考simple_GROP项目训练的模型,GSM8K,训练了200个step,出现了一次however。 使用了3张A800 80G,训练了20多分钟
训练结果:
loss
GPU
memory
测试结果
demo_math_chat_gen(simple_GRPO_why)
demo_math_chat_gen(Qwen2.5-7B)
notice
在GSM8K上进行评估,Qwen2.5-7B的得分为85.4。原因可能是是
https://github.com/open-compass/opencompass/issues/1878