SGOCR is a spatially-grounded OCR visual question answering dataset for training and evaluating models that must read, localize, and reason about text in images.
The dataset contains grounded question-answer pairs over ChartQA, TextOCR, and COCO/COCO-Text source images. It is designed for OCR-aware VQA, text grounding, region-conditioned QA, and data-centric experiments around scene text understanding.
Project repository:… See the full description on the dataset page:
https://huggingface.co/datasets/dreeseaw/SGOCR.