AlphaNeural
grpo_Qwen-Qwen2.5-7B-Instruct_ref_math_easy_Qwen-Qwen2.5-1.5B-Instruct_kl1.0_lr1e-6 – AI Model by ChenWu98 | AlphaNeural AI