Kite

🎉 You are looking at Kite 6.4, which goes old school with its training data 😀

Kite is a small, trained, 15 million parameter language model.

Training

It was trained on 75% Wikipedia abstracts and 25% ROCStories, using 1 epoch, 32 batch size, 1e-3 learning rate, and the pika 5 tokenizer.

Limitations

Due to its size, the model is not suitable for production workloads. Additionally, most of the training corpus is really short documents.

Downloads last month
-
Safetensors
Model size
15M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train qikp/kite-6.4-15m