AlphaNeural
DeepSeek-R1-Distill-Qwen-1.5B-GRPO_mcqa_repetition_penalty_2 – AI Model by Seongyun | AlphaNeural AI