Continued from the 4.3 lineage with Gemma staying securely in her role as judge rather than rewriting directly.
600+ steps RLVR GRPO in spicy roleplaying first-turns, with Gemma 4 31b as judge paying attention to writing craft;
followed by DPO including self-rewrites of about half of the particularly bad trajectories from RLVR
(until Gemma rated them well in comparison, and particular issues like a godmoded POV were fixed).
Ran 5 seeds at low batch size to explore and merged via Karcher mean.
A mix of approaches including utilizing and not-utilizing the <feelings> tags were tried in RLVR.
She's trying some interesting things with the tokens, including compound-word emotions that split the difference between word and phrase,
but the ratings aren't notably better with the tags than without them -
so I wouldn't consider her a reasoning model yet, and it's not in the chat template.
More Arsenic self-portraits:
image
image
Amethyst:
image
image
image
Aegis:
image
)
image
image
arsenic-v4.4-merged
This is a merge of pre-trained language models created using mergekit.
Merge Details
Merge Method
This model was merged using the Karcher Mean merge method.