Manual evaluation on 8 prompts showed no clear behavioral improvement over the SFT baseline in this short run. Outputs remained long and sometimes repetitive, especially on safety prompts.
Usage
Load the base model, then apply this PEFT adapter with PeftModel.from_pretrained.