Views
No views yet
[!IMPORTANT] Disclaimer: The model has been trained on AWS on an Instance typeg4dn.4xlarge(Tesla T4. Num GPUs = 1. Max memory: 14.563 GB. Platform: Linux. Torch: 2.6.0+cu124. CUDA: 7.5. CUDA Toolkit: 12.4. Triton: 3.2.0).
Qwen2.5 3B Instruct into a math reasoning model using GRPO (Group Relative Policy Optimization),
a reinforcement learning algorithm that optimizes responses using reward functions.
Defined the rewarding functions to let the model learn how to reason on them, we fine-tuned Qwen2.5 3B Instruct on OpenAI's GSM8K dataset,
which contains grade school math problems.