This dataset is a bucketed WebDataset-style export of
Yesianrohn/OCR-Data,
with images paired with concise English captions for text-to-image training.
The source dataset aggregates public OCR benchmarks with images, recognized
text, text-region bounding boxes, and polygon annotations. This export keeps the
source OCR metadata in the JSON sidecars and replaces the training captions with
short descriptions focused on the visible image content, readable… See the full description on the dataset page:
https://huggingface.co/datasets/data-archetype/ocr_captions.