Golden paired dataset for training models to transliterate Arabic Latin text into
scholarly diacritized form — built from a single recorded Islamic lecture
(Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly
transcript.
Two artifacts are stored separately for provenance and review:
bronze.jsonl
773
Bronze — every aligned sentence pair… See the full description on the dataset page:
https://huggingface.co/datasets/olanigan/logical-transcripts.