AlphaNeural
grpo_thinking_ultrafeedback-original_32_64_4_3e-3_2e-7_step-120_1.7B – AI Model by RLAIF | AlphaNeural AI