This was mostly a test to see what the loss/eval looked like when training on top of Harmonia, and in that sense it was a sterling success, without the "jitter" I experienced training on top of Nethena 20b.
Quick testing shows a bit of derpiness, but a nice conversational flow. Overall, this will be helpful in developing additional 20b merges.
This model is a fine-tuned version of
athirdpath/Harmonia-20B on the HF No Robots dataset.
It achieves the following results on the evaluation set: