🎉 You are looking at Kite 6.1 Release 1, which is our 3rd attempt at combatting underfitting!
Kite is a small, trained, 18 million parameter language model.
It was trained on a tokenized version of
qikp/small-data-3.2, which is a mixture of various datasets, using 2 epochs, 32 batch size, 1e-3 learning rate, and the
pika 4 tokenizer.
Also, evaluation on a tokenized and truncated
byunggill/gpt-2-output was done during training.
Due to its size, the model is not suitable for production workloads.