This model is a fine-tuned version of Meta's Llama 3.1 8B model that has been trained to improve its own abilities through reinforcement learning. The model has been trained using GRPO (Generalized Reinforcement from Preference Optimization) to better follow instructions and generate high-quality responses.
Training Procedure
The model was trained using Unsloth's fast training framework with the following key parameters: