Source dataset: docling-project/ibm-handwriting-campaign
This directory contains a Hugging Face dataset export generated from the project source handwriting data.
Splits: train, validation, test
Total samples: 841 (train: 671, validation: 83, test: 87)
Format: Parquet files with images stored as bytes
Each record corresponds to a single word extracted from scanned forms and includes the word image, annotation JSON… See the full description on the dataset page:
https://huggingface.co/datasets/Felix92/ibm-handwriting-campaign-word.