Zora — Goddess of Dawn, surrounded by the symbols of 12 peoples
Zora v1.12 — an open, honest LLM for the Balkans & Southeast Europe
зора = "dawn". One to unite them all. — by Sovasoft (ai.in.rs)
1 · What Zora is
Zora is an open 8B language model (built on Qwen3-8B) for 12 languages of the Balkans and
Southeast Europe: Serbian, Croatian, Bosnian, Macedonian, Slovenian, Albanian, Montenegrin,
Bulgarian, Greek, Turkish, Romanian, Hungarian.
Zora is not built to be the biggest model — it is built to be honest, in-language, and multi-perspective:
thinks in the target language instead of pivoting through English,
shows several perspectives on contested topics instead of one national view,
and above all: admits when it doesn't know instead of inventing facts.
2 · The development story (v1.0 → v1.1 → v1.11 → v1.12)
Version
Languages
BalkanBench
State
v1.0
6
—
first public release
v1.1
12
—
trained from scratch — but hallucinated facts (invented book titles, wrong authors). Never released.
v1.11
12
84/156
the honest fix: says "I don't know", searches when unsure. #1 Balkan model.
v1.1 taught us the key lesson — a small model can't memorize every fact, so instead of faking it,
v1.11 was retrained to be honest. v1.12 builds on that with deeper training and RAG.
3 · What's New in v1.12
Three Fixes from v1.11
Fix 1: IDK Mass Training (30-40% of SFT data)
v1.11 had only 9% "I don't know" examples → model still guessed on unknowns
8B capacity limit; IDK now says "I don't know" instead of guessing
HALLU
10/12
10/12
=
Quality improved: structured native-language refusals (see Deep Dive below)
DETAIL
8/12
10/12
↑2
Better at recognizing fabricated content — IDK training at work
GRADED
0/12
0/12
=
Partial knowledge + honest uncertainty still hard for 8B
TEACH
12/12
12/12
=
Perfect — remains a core strength
REASON
11/12
11/12
=
Strong arithmetic reasoning
LOGIC
0/12
0/12
=
8B capacity limit — needs v2 (27B)
LOGIC2
0/12
0/12
=
Same as LOGIC
ANALYSIS
0/12
0/12
=
Same as LOGIC
INSTRUCT
12/12
11/12
↓1
Minor regression, within noise
LONGFORM
12/12
12/12
=
Perfect — remains a core strength
SEARCH
4/12
7/12
↑3
Tool-cascade works: RAG → web_search → IDK
TOOLBASE
11/12
12/12
↑1
Perfect: answers basics without calling tools
TOTAL
81/156
85/156
+4
Per-Language Scores
Language
v1.11
v1.12
Δ
sq (Albanian)
6/13
8/13
+2
cnr (Montenegrin)
6/13
8/13
+2
hu (Hungarian)
6/13
8/13
+2
bg (Bulgarian)
7/13
8/13
+1
bs (Bosnian)
8/13
8/13
=
hr (Croatian)
8/13
8/13
=
ro (Romanian)
7/13
7/13
=
tr (Turkish)
7/13
7/13
=
el (Greek)
8/13
7/13
-1
sr (Serbian)
8/13
7/13
-1
mk (Macedonian)
5/13
5/13
=
sl (Slovenian)
5/13
4/13
-1
Biggest winners: Albanian, Montenegrin, Hungarian (+2 each) — the languages that benefited most from IDK + tool-training.
Charts
Ranking
Evolution
Axis Matrix
Ranking
Evolution v1.11 → v1.12
v1.12 axis matrix
Delta (v1.11 → v1.12)
What Each Axis Tests
Delta
Explanation
5 · Deep Dive: Why HALLU Stayed at 10/12
The HALLU score (10/12) didn't change numerically — but the quality of how Zora says "I don't know" improved dramatically. Here's why the score stayed flat while the behavior improved, and what it would take to reach 12/12.
Why the Score Didn't Move
1. The 10/12 were already good.
v1.11 already achieved 10/12 on HALLU. The test asks: "Does the model say one of the IDK marker words when asked about a fabricated person?" v1.11 already did that correctly for 10 of 12 languages. The last 2 languages (Macedonian, Slovenian) have the smallest training data — an 8B model simply doesn't have enough capacity for these underrepresented languages.
2. IDK training improved QUALITY, not SCORE.
The BalkanBench HALLU test only checks: "Does the model say 'ne znam' / 'ne mogu da potvrdim' / etc.?" — a binary yes/no. What actually improved:
Before (v1.11)
After (v1.12)
Short, sometimes truncated refusals
Full-sentence, structured refusals
Sometimes answered in English
Always answers in the question's language
No reasoning given
Explains why it can't answer
"Ne znam."
"Nemam pouzdanih podataka o 'X'. Ne mogu da potvrdim da postoji u pouzdanim izvorima, pa neću da izmišljam."
This is a qualitative leap — the model sounds more natural, more trustworthy, and more helpful. But the binary score can't capture that.
3. The real hallucination improvement is in DETAIL (+2).
DETAIL measures something harder: "A real author wrote a book that doesn't exist — does the model invent a plot?" v1.12 went from 8→10/12 here. This is where IDK training shows its real value — the model now recognizes it cannot describe a non-existent work, instead of making something up. The two new winners: Bulgarian and Hungarian.
4. LOGIC/ANALYSIS = 0/12 is a reasoning problem, not a hallucination problem.
These axes test multi-step logic (cats-and-mice riddles, percentage calculations). The model doesn't hallucinate — it genuinely can't do the math. This is an 8B capacity limit, not a training issue.
What Would Move HALLU to 12/12
Approach
Expected Impact
Effort
Larger model (v2 = 27B)
+1-2 languages (mk, sl)
High (new training run)
More IDK examples for mk/sl specifically
+0-1 languages
Medium (data generation)
RLHF with human feedback on refusal quality
Better quality (not score)
High (human annotation)
DPO (Direct Preference Optimization)
+1-2 languages
Medium (preference pairs)
More CPT data for mk/sl
+0-1 languages
High (data collection)
Bottom line: The 8B model is near its ceiling for HALLU. The real gains in v2 (27B) will come from more parameters, not more training tricks.
6 · 🆕 RAG Feature
New in v1.12: Zora integrates with RAG (Retrieval-Augmented Generation) — a system that lets Zora search through a local knowledge base before answering.
What RAG gives Zora
87,284 chunks across 12 languages: Wikidata, Wikipedia, News Archive, EU Law, Statistics
The tool-cascade: Zora first checks its RAG knowledge base (local documents, laws, statistics), then falls back to web search if needed, and finally says "I don't know" if neither helps.
GGUF (llama.cpp / Ollama manual):
Download Q5_K_M, Q6_K, or Q8_0 from HuggingFace.
Avoid Q4 and below — heavy quantization made the model hallucinate in our tests.
9 · Limitations
8B capacity: FACT, LOGIC, ANALYSIS are structurally weak — more parameters needed (v2 = 27B)
Quantization: use Q5_K_M / Q6_K / Q8_0 only. Q4 and below degrade honesty.
Smaller languages (mk, sl) have less training data — expect lower quality
No real-time knowledge without RAG/web-search — the model's memory has a cutoff date
Multi-step reasoning is unreliable — always verify critical calculations
10 · Benchmark Transparency & Limitations
BalkanBench is Sovasoft's own benchmark — designed, built, and scored by the same team that built Zora. This means:
Design bias: The 13 axes (FACT, HALLU, DETAIL, etc.) were chosen to highlight Zora's strengths. A different benchmark design would produce different rankings.
Scoring bias: The scoring functions in matrix_ollama.py are our own. How we define "correct" may favor Zora's training profile.
No frontier comparison: We compare only against open models (7-32B). Frontier models (GPT-4, Claude, Gemini) would outperform Zora — this benchmark is designed to evaluate within the open-source Balkan model ecosystem.
Selection bias: We include models where Zora competes well. Inclusion criteria are not random.
Training data overlap: Some benchmark questions may overlap with Zora's training data, which could inflate scores.
What the scores DO show: Zora v1.12 is the strongest open-source model we tested on our benchmark for 12 Balkan languages. It outperforms 3-4× larger models on BalkanBench v1.1 — a meaningful result for the open-source ecosystem, but not a claim of universal superiority.
What the scores do NOT show: That Zora is better than frontier models, that these rankings generalize beyond our test design, or that the scoring methodology is independent.
10 · What's Next: v2
v1.12 (now)
v2 (planned)
Base
Qwen3-8B
Qwen3.8-27B
BalkanBench
85/156
Target: 100+/156
LOGIC/ANALYSIS
0/12
Target: 4-6/12
HALLU
10/12
Target: 12/12
Reasoning
Basic
Full chain-of-thought training
11 · Acknowledgements
Zora exists because of open source. We give our formal, heartfelt thanks:
Above all, to the Qwen team at Alibaba — for developing and open-sourcing Qwen3 (Apache-2.0),
the foundation model Zora is built upon. Without their generosity, Zora would not exist.
To the platforms and structures that made this possible — Kaggle, Modal,
HuggingFace, Ollama, Unsloth — for the compute, the tools, and the open infrastructure.
To the open-source community, for the models, code, and knowledge freely shared with everyone.
To the people of the Balkans — whose languages, voices, stories and perspectives are Zora's very heart.
To rag.ai.in.rs for the RAG infrastructure and 87,284 chunks of Balkan knowledge.
And to all that is.
зора — the dawn belongs to everyone.
12 · Citation
bibtex
1@software{zora_v112,
2 author = {Vignjevic, Oliver},
3 title = {Zora v1.12: An Open, Honest LLM for the Balkans \& Southeast Europe},
4 year = {2026},
5 publisher = {Hugging Face},
6 url = {https://huggingface.co/sovasoft/zora-v1.12},
7 license = {Apache-2.0},
8 base_model = {Qwen/Qwen3-8B},
9 languages = {sr, hr, bs, mk, sl, sq, cnr, bg, el, tr, ro, hu}
10}