Finetuned Qwen1.5B (R1 Distilled Version) on
this dataset,
which comes from
this dataset but with an
additional "summary" produced by an in-house synthetic data generator.
This LoRA is therefore a LoRA which helps the model return a "\n\nFinal Answer: ..." after it's reasoning and initial response steps.
This qwen2 model was trained 2x faster with
Unsloth and Huggingface's TRL library.