LoRA adapter further fine-tuned via Direct Preference Optimisation (DPO)
on top of the SFT model. Trained to generate more natural conversation
endings and diverse NPC responses.
This adapter is applied on top of the merged SFT model, not directly
on the base Llama model. See Usage section.