Kite is a small, trained, 12 million parameter language model.
It was trained on 250K selected passages of MS MARCO, using 1 epoch, 32 batch size, 1e-3 learning rate, and the
MicroSupra tokenizer.
Due to its size, the model is not suitable for production workloads.