A bunch of datasets preprocessed and formatted with
https://github.com/openai/openai-python/blob/main/chatml.md (with an addition of a context message to help RWKV (no lookback))
The dataset makes use of two more tokens. You will need to use the supplied 20b_tokeniser file with both training and inference.
English mainly, might be a few bits of other languages.
Improve system prompt effect on output.
Get more reasoning… See the full description on the dataset page:
https://huggingface.co/datasets/m8than/raccoon_instruct_conversation.