tl;dr: This model demonstrates a critical lesson in language modeling: loss ↓ ≠ factual accuracy ↑. Despite achieving PPL 42.7, it confidently generates wrong facts. This repo documents why.
🎯 The Core Problem
After much training steps on 500K Turkish documents:
Metric
Start
End
Improvement
Validation Loss
6.0
3.75
37% better ✅
Validation PPL
397
42.7
90% better ✅
Factual Accuracy
❌
❌
Still inconsistent
📉 Loss vs Factuality Divergence
Training Progression for prompt "Türkiye'nin başkenti"
Step
Val Loss
Val PPL
Generated Capital
Correct?
1000
5.98
397.3
Ankara
✅
3000
3.94
51.7
Ankara
✅
5000
4.02
56.2
Random city
❌
6500
3.90
49.6
Bolu
❌
7500
3.83
46.1
Konya
❌
8000
3.80
44.8
Bursa
❌
9000
3.75
42.7
Ankara (sometimes)
⚠️
Key observation: Loss steadily decreases, but capital city prediction remains unstable.
🧪 Concrete Examples
Prompt: "Türkiye'nin başkenti"
Step 6500 output:
"Türkiye'nin başkenti Bolu'da 24 Mayıs 2012 tarihinde toplanan ve 23. dönem TBMM'de kabul edilen 6360 sayılı kanun..."
❌ Wrong: Bolu is not the capital
✅ Right: Date format, legal language, formal tone, grammar
Step 7500 output:
"Türkiye'nin başkenti Konya, Türkiye'nin en büyük kenti Konya'dan Avrupa'ya yayılan ilk yatırımını 25-30 Mayıs 2020..."
❌ Wrong: Konya is not the capital
✅ Right: Geographic context, date ranges, economic terminology
Step 9000 output:
"Türkiye'nin başkenti Ankara'da düzenlenen Dünya Kadınlar Basketbol Şampiyonası'nda..."
✅ Finally correct!
🤔 Why This Happens
What the Model Actually Learns
Cross-entropy loss optimizes for: "What token is likely in this context?"
In training data distribution:
"Türkiye'nin başkenti Ankara..." appears ~60% of patterns
"Başkent Bursa/Konya/İzmir..." appears ~40% (from various contexts)
The model learns distributional probabilities, not factual truth.
From the model's perspective:
Sometimes generate "Ankara" (most frequent)
Sometimes generate other cities (contextually plausible)
Both reduce loss equally if they appear in training data
Why Loss Still Decreases
Even with wrong facts, the model improves at:
✅ Grammar (Turkish morphology)
✅ Syntax (sentence structure)
✅ Style (formal/informal tone matching)
✅ Context coherence (topic consistency)
✅ Pattern matching (Wikipedia-style text)
Loss measures linguistic fluency, NOT factual correctness.
@misc{kayra2024hallucination,
title={Why Small Turkish GPTs Hallucinate Facts: An Experimental 85M Model},
author={sixfingerdev},
year={2024},
publisher={HuggingFace},
howpublished={\url{https://huggingface.co/sixfingerdev/kayra-1-exp}},
note={Research on loss-factuality divergence in low-resource language models}
}
🙏 Acknowledgments
Inspiration: Eleuther AI's research on small model limitations
Data: Wikimedia Foundation, Common Crawl (mC4)
Framework: PyTorch, HuggingFace Transformers
📜 License
MIT License - Use freely for research and education.
Disclaimer: This model is intentionally shared with its flaws documented. It serves as a learning resource demonstrating why small LMs hallucinate, not as a production tool.
Kayra-1-exp - Teaching us what 85M parameters cannot do 🔬
Discussion: Found interesting hallucination patterns? Share your findings in the community discussions tab. Let's learn together why small LMs hallucinate. 🇹🇷
🌙 Kayra-1-exp
Kayra - Sıfırdan Türkçe ile eğitilmiş ilk deneysel GPT modeli.
📊 Model Detayları
Model türü: Decoder-only Transformer (GPT-style)
Parametreler: ~85 milyon
Validation PPL: 42.7
Validation Loss: 3.75
Dil: Tamamen Türkçe
Lisans: MIT
🏗️ Mimari
Layers: 10
Hidden size: 640
Attention heads: 10
FFN size: 2560
Vocabulary: 32,000
Context length: 512
📚 Eğitim Verisi
Wikipedia TR: ~170K makale
mC4 Turkish: ~330K doküman
Toplam: ~500K dedupe edilmiş doküman (MinHash LSH)