A fairness-enhanced mini subset of the IAM (Institute of Applied Mathematics) Handwriting Database with 500 randomly selected samples. Each image is separated into printed and handwritten components for fair VLM evaluation.
The original IAM dataset contains images with both printed reference text and handwritten content on the same page. This allows Vision Language Models (VLMs) to "cheat" by… See the full description on the dataset page:
https://huggingface.co/datasets/kenza-ily/iam_disco.