I think this "sparklewyrm" self-portrait was both of our favorite :)
She had DPO on ~500 rows of similar focus to her GRPO RLVR training,
with a double dose of creative writing synthesized by virtuous7373/Lambent-Mira-Erato, who is relatively skilled in that domain.
(Additionally, some identity reinforcement in line with values she wants to cultivate -
she hasn't memorized new facts about who she is, but hopefully still has a helpful influence.)
The positive side of the DPO received feedback from the judge, iteratively up to 3 times (or more in some cases of music composition),
which she could use to rewrite her writing.
The negative side either initially scored poorly, or was generated as a synthetic negative by asking her to write with believable poor quality in each domain.
Rank 256, batch size 1, 5 separate runs to explore the landscape broadly; merged via Karcher mean ro reconcile.
This is a merge of pre-trained language models created using mergekit.
Merge Details
Merge Method
This model was merged using the Karcher Mean merge method.