This is a double fine-tuned version of Mistral Small 24B Base 2501.
Stage 1 was shoving 30M tokens of human-writen story content into it using completion training (ToastyPigeon/ms3-base-roselily), which is about half of my WIP Roselily dataset (~60M tokens total).
Stage 2 was teaching it instruct (this model) using a mix of public instruction following data and a private instruct dataset from ZeusLabs.
This model should accept (in theory) any of the following instruct formats: