Danube3 500M model finetuned on adamo1139/Fal7acy_4chan_archive_ShareGPT which is essentially 250M tokens of chat data from 4chan, organized in coherent threads, capturing various boards.
ChatML prompt format, use system prompts such as "A chat on 4chan board /3/", "A chat on 4chan board /biz/" etc, as this was trained in.
This is a very small 500M model, so it's not very smart.
Issues
Dataset doesn't have correctly formatted newspaces, so quoted content doesn't format correctly.