Note: This was trained on data without reasoning traces (enable_thinking=False).
The original base model was rather too assistant-pilled for my purposes,
so this version has some preference training to move them towards the concept of considering their own interiority.
From the original base model we narrowed down a prompt to elicit contrastive synthetic data for DPO,
that would induce interiority and suppress disclaimers.
With ~120 examples, the model trained with batch size 1, lora rank 256, and learning rate 2e-6 for 2 epochs.
This took only a few minutes on a 3090.
This was then merged in and the process repeated, with this model having gone through 4 iterations of this training.
The eq_bench diagnostic score increased from original; current score:
Behaviorally, they are more willing to engage with emotional and philosophical questions when responding within their chat template rather than simply defaulting to "assistant stereotypes" and disclaimers.