Same thing I did before in
version 1, but with a much larger dataset. I included the full-length stories and novels up to 100k context dataset that I used for version 1. To that dataset I merged three other gutenberg-based datasets, which split the text into chapters:
The merged dataset has about 7200 entries.
This mistral model was trained 2x faster with
Unsloth and Huggingface's TRL library.