AlphaNeural
GRPO-edu_feedback-train16-n8-cosine-Qwen-Qwen2.5-3B-Instruct-biochem-mix20-3b-lr5e-6-maxstep100 – AI Model by Johnny1024 | AlphaNeural AI