A computational study of how similar Naga languages are to each other, measured using parallel Bible translations as a controlled corpus. 36 Naga and Naga-adjacent languages plus English — compared using character-level patterns to find shared vocabulary, cognates, and structural connections.
The goal is not to rank languages or tribes. It is to find the common threads — the shared words, the similar sounds, the patterns that connect communities across mountains and state borders.
Key Findings
1. Six similarity methods evaluated; three are primary
Method
Cohen's d (Full)
ARI (Full)
Cohen's d (NT)
ARI (NT)
Status
Jaccard (top-1000 trigrams)
1.64
0.65
1.57
0.55
✅ Best clustering
TF-IDF (2-4 char n-grams)
1.59
0.46
1.94
0.66
✅ Best pair ranking
Subword Vocabulary JSD
1.27
0.44
1.37
0.32
✅ Complement
Character Trigram Frequency
1.11
0.36
1.12
0.35
⚠️ Marginal
Character Frequency
0.63
0.28
0.63
0.29
❌ Fails
Glot500-m Embeddings
0.27
0.15
0.27
0.15
❌ Fails
2. Known language families cluster correctly
Dendrogram — Jaccard
The dendrograms, t-SNE, and PCA all confirm:
Ao (Central Naga) and South Patkaian cluster together
Angami-Pochuri forms a distinct cluster
Monsang-Moyon pair is the most similar in the dataset
English is correctly the most distant language
3. Word extraction reveals shared vocabulary
Using NPMI (Normalized PMI) on verse-aligned parallel text + independent RAT (TF-IDF + BM25 retrieval):
~2,000 English concepts mapped to 36 Naga languages