Model-inferred diacritization (tashkeel) of 6,934,210 Arabic poetry verses
from the arbml/ashaar dataset.
Raw verses: arbml/ashaar (241,964 poems, ~6.98M verse lines).
Diacritization model: basharalrfooh/Fine-Tashkeel
(ByT5-Large, fine-tuned on classical Arabic Tashkeela corpus).
Inference settings: FP16 on a single NVIDIA RTX 5880 Ada (48 GB),
max_new_tokens=48, max_input_len=128, greedy decoding. ~19 hours end-to-end
at ~100… See the full description on the dataset page:
https://huggingface.co/datasets/mysamai/ashaar-tashkeel.