OCR training data from AIDA-project
Dataset Summary
The zip file contains textlines and their annotations from AIDA-project. There are ~ 166k textlines that are mainly in Finnish language, but contain a little Swedish and
English and little French and German textlines. The textlines contains typewritten and also handwritten lines. Roughly 24 % of the annotated lines are handwritten and the rest are
typewritten. The dataset also contains 120 000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Kansallisarkisto/AIDA_ocr_training_data.