HAFITH: Aspect-Ratio Preserving Vision-Language Model for Historical Arabic Manuscript Recognition
State-of-the-art OCR model for historical Arabic manuscripts achieving 5.10% CER through native-resolution encoding, Arabic-native tokenization, and synthetic pretraining.
Model Summary
Architecture: Vision-Language (Encoder-Decoder)
Vision Encoder: SigLIP V2 NaFlex (400M params, preserves aspect ratios up to 20:1)
Text Decoder: RoBERTa-Large (242M params, trained from scratch)
Operates on pre-segmented text lines (requires line segmentation for full pages)
Trained on modern Arabic vocabulary (may miss some archaic terms)
Performance degrades on severely damaged manuscripts (>9% CER)
Maximum line length limited by 512-patch budget
Comparison with Baselines
Model
Encoder
Tokenizer
CER
WER
CRNN+CTC
CNN
Character-level
14.82%
-
TrOCR-Base
BEiT-B (384×384)
RoBERTa
13.41%
-
TrOCR-Large
BEiT-L (384×384)
RoBERTa
11.73%
31.82%
HATFormer
BEiT-L (384×384)
RoBERTa
8.60%
-
HAFITH (Ours)
SigLIP2 NaFlex
Aranizer
5.10%
18.05%
Citation
bibtex
1@article{naseif2026hafith,
2 title={HAFITH: Aspect-Ratio Preserving Vision-Language Model for Historical Arabic Manuscript Recognition},
3 author={Naseif, Mohammed and Mesabah, Islam and Hajjaj, Dalia and Hassan, Abdulrahman and Elhayek, Ahmed and Koubaa, Anis},
4 journal={arXiv preprint arXiv:XXXX.XXXXX},
5 year={2026}
6}