Finetuned from model: unsloth/Llama-3.2-3B-Instruct
This Llama 3.2 3B model was fine-tuned using Group Relative Policy Optimization (GRPO) on the GSM8K dataset for improved mathematical reasoning capabilities. It was trained 2x faster with Unsloth and Huggingface's TRL library.