English only. Trained exclusively on English text.
No instruction tuning. Base language model only, not aligned for chat or task completion.
Partially trained. This checkpoint represents 6,000 of 50,000 planned steps (~12% of full training). Quality will improve significantly with full training.
Limited capacity. At 140M parameters, the model underperforms larger models on knowledge-intensive benchmarks. Outputs may be factually incorrect.
Context window. 2,048-token context window.
No formal evaluation. Benchmark results have not been reported for this release.
Citation
bibtex
1@software{kilat2026,
2 author = {Abdul Wahid Rukua},
3 title = {Kilat: Kernelized Lightweight Transformer Training Framework},
4 year = {2026},
5 url = {https://github.com/Airukua/kilat}
6}
78@misc{babykilat2026,
9 author = {Abdul Wahid Rukua},
10 title = {BabyKilat: A Lightweight MoE Language Model},
11 year = {2026},
12 publisher = {Hugging Face},
13 url = {https://huggingface.co/AiRukua/BabyKilat}
14}