This is artificial Faroese OCR training data created by collection real OCR errors and inserting them into 38 million tokens of non-OCRed text.
The parallel data is set up as a TSV file with the first column (fo_err) being the text with OCR errors, while the second column (fo_corr) is without OCR errors.
This dataset was created by using scripts from
https://github.com/atlijas/ocr-post-processing.
Two ByT5 models have been fine-tuned with the data:… See the full description on the dataset page:
https://huggingface.co/datasets/AnnikaSimonsen/FO-OCRtrain.