Based on alpindale/Llama-3.2-3B-Instruct
And trained with dataset openai/gsm8k
The objective of this model is to test the novel GRPO training used in DeepSeek R1.
Using the reinforcement learning (RL) algorithm to improve the reasoning capabilities of the Llama-3.2-3B-Instruct.