The smartest L3 8B model combined with high-end RP model. What could go wrong.
The idea was to fuse a bit of SimPO's realism with Stheno. It took a few days to come up with a balanced slerp configuration, but I'm more than satisfied with the end result.
Changes compared to v3.1
- Included a mix of SFW and NSFW Storywriting Data, thanks to Gryphe - Included More Instruct / Assistant-Style Data
- Further cleaned up Roleplaying Samples from c2 Logs -> A few terrible, really bad samples escaped heavy filtering. Manual pass fixed it.
- Hyperparameter tinkering for training, resulting in lower loss levels.
Testing Notes - Compared to v3.1
- Handles SFW / NSFW seperately better. Not as overly excessive with NSFW now. Kinda balanced.
- Better at Storywriting / Narration.
- Better at Assistant-type Tasks.
- Better Multi-Turn Coherency -> Reduced Issues?
- Slightly less creative? A worthy tradeoff. Still creative.
- Better prompt / instruction adherence.