This dataset contains segmented line-level images and corresponding transcriptions in Luganda, a low-resource Bantu language spoken primarily in Uganda. It was created to support research in optical character recognition (OCR) and handwritten/printed text recognition for under-resourced African languages.
The dataset was developed as part of an OCR research project focused on building end-to-end deep learning models for text… See the full description on the dataset page: https://huggingface.co/datasets/marconilabmak/luganda-ocr-dataset.