This repository contains a synthetic OCR dataset generated by the Synthetic OCR Image Generator pipeline.
It is intended for benchmarking and evaluating OCR or VLM systems on markdown-oriented document recognition tasks.
Dataset:
https://huggingface.co/datasets/junyeong-nero/synthetic-ocr-images-ko
Language: ko
Generated samples: 1000
Requested samples: 1000
Generation mode: markdown