AlphaNeural
grpo_rl_ref_math_hard_Qwen-Qwen2.5-Math-1.5B-Instruct_rlglobal_step_3400_kl2.0_lr1e-6 – AI Model by ChenWu98 | AlphaNeural AI