KoshurAI_Tarjuma_v2 — Kashmiri Continual Pretraining Base
⚠️ This is a base language model, not a translation model.
For Kashmiri ↔ English translation, use the fine-tuned adapter:
Faizaniqbal/KoshurAI_Tarjuma_v3
What is this?
KoshurAI_Tarjuma_v2 is a 5B-parameter Gemma 3 model that has been
continually pretrained on 2.8 million tokens of native Kashmiri text,
giving it deep knowledge of the Kashmiri language (Perso-Arabic script).
It serves as the Stage 1 base in the KoshurAI two-stage pipeline:
Omarrran/koshur-kouter-ks-en_v1 ← fine-tuned on Kashmiri–English pairs
↓ continual pretraining on 2.8M Kashmiri tokens
Faizaniqbal/KoshurAI_Tarjuma_v2 ← this model (language knowledge)
↓ SFT on 16,637 EN↔KS pairs (LoRA adapter)
Faizaniqbal/KoshurAI_Tarjuma_v3 ← final translation model
Model Details
|
|
|
|
|
|
|
**
Author
**
|
Faizan Iqbal (
@Faizaniqbal
)
|
|
**
Base model
**
|
Omarrran/koshur-kouter-ks-en_v1
|
|
**
Architecture
**
|
Gemma3ForCausalLM (5B parameters)
|
|
**
Tensor type
**
|
BF16
|
|
**
Pretraining data
**
|
2.8M tokens of Kashmiri text
|
|
**
Languages
**
|
Kashmiri (ks · kas_Arab), English (en)
|
|
**
License
**
|
Apache-2.0
|
Pretraining Corpus
The 2.8M token corpus was assembled from two sources:
InPage documents — professionally published Kashmiri literature,
journalism, academic scholarship, and religious texts spanning multiple
decades, converted to Unicode via a custom InPage converter.
Native speaker translations — texts translated into Kashmiri by
native speakers, providing natural human-authored language coverage.