Kite is a small, trained, 15 million parameter language model.
It was trained on 75% Wikipedia abstracts and 25% ROCStories, using 1 epoch, 32 batch size, 1e-3 learning rate, and the
pika 5 tokenizer.
Due to its size, the model is not suitable for production workloads. Additionally, most of the training corpus is really short documents.