FIN-13K is a Finnish OCR dataset containing 13,037 image-text pairs extracted from historic Finnish newspapers and journals. The dataset is derived from the FIN-BERT dataset and is suitable for training and evaluating OCR models on Finnish text.
Optical Character Recognition (OCR): Recognizing text from line-level images
OCR Post-correction: For… See the full description on the dataset page:
https://huggingface.co/datasets/caveman273/fin-13k.