This is the GGUF conversion of the Llama 3.2 3B model fine-tuned using Group Relative Policy Optimization (GRPO) on the GSM8K dataset for improved mathematical reasoning capabilities. The finetuning was performed 2x faster with
Unsloth and Huggingface's TRL library.
An Ollama Modelfile is included for easy deployment.