This dataset contains synthetic line images meant for fitting OCR models for North, South, Lule and Inari Sámi.
Clean line images are created using Pillow and they are subsequently distorted using Augraphy [1].
The text in this dataset comes from Giellatekno's corpus. Specifically, we used the data files of the converted/-directories of [2][3][4][5] (commit hashes… See the full description on the dataset page:
https://huggingface.co/datasets/magwrap/synthetic_sami_ocr_data.