Had this one in the works for a while, but was struggling to find the right hyperparams to get this model to behave nicely. Thank you to TheDrummer for helping me out with them.
This model is a creative writing and RP model. It's pretty verbose. The intent is to keep the behavior of the original model, but to slightly improve writing, dialogue & creativity.
SFT on approx 10 million tokens, SFW / NSFW RP, stories, creative instruct & chat data.
MoE are brutal to train even with a small dataset like mine, so I took a different approach from usual. I used a very low LR in an effort to avoid having to apply DPO / KTO training afterwards.
I think there's likely a better config to be found, but experimentation with the model to find it is quite draining.
>
Axolotl configs
Not optimized for cost / performance efficiency, YMMV.