Views
No views yet
| Property | Value |
|---|---|
| Base Model | unsloth/DeepSeek-OCR |
| Model Type | Vision-Language Model (VLM) |
| Fine-tuning Method | LoRA (Low-Rank Adaptation) |
| Language | Sinhala (සිංහල) |
| License | Apache 2.0 |
| Parameters | ~3.5B (base) + 155M (LoRA trainable) |
| Precision | 4-bit quantized (inference) |
| Metric | Score | Description |
|---|---|---|
| Character Accuracy | 98.0% | Percentage of correctly recognized characters |
| CER (Character Error Rate) | 0.020 | Lower is better (0 = perfect) |
| WER (Word Error Rate) | 0.045 | Word-level accuracy |
| BLEU Score | 0.965 | Text similarity score (0-1) |
| ANLS | 0.980 | Average Normalized Levenshtein Similarity |
| METEOR | 0.975 | Semantic similarity score |
| Accuracy Range | Number of Samples | Percentage |
|---|---|---|
| ≥ 99% | 65/202 | 32.2% |
| ≥ 95% | 145/202 | 71.8% |
| ≥ 90% | 185/202 | 91.6% |
| ≥ 80% | 197/202 | 97.5% |
| < 80% | 5/202 | 2.5% |
| Model | Character Accuracy | CER | Training Samples |
|---|---|---|---|
| This Model (A100, 6 epochs) | 98.0% | 0.020 | 707 |
| Baseline (P100, 3 epochs) | 96.98% | 0.030 | 707 |
| Improvement | +1.02% | -33% | - |
## 🎓 Citation
If you use this model, dataset, or research in your work, please cite our paper:
```bibtex
@misc{dilhara2026crosstemporalsinhalaocrpagelevel,
title={Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis},
author={Avisha Dilhara and Nevidu Jayatilleke},
year={2026},
eprint={2606.29378},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={[https://arxiv.org/abs/2606.29378](https://arxiv.org/abs/2606.29378)}
}