Synthetic English OCR Detection and Recognition 240K
π Current dataset size: 240,000 paired OCR samples
The current v2.0 release contains exactly 240,000 detector images and
240,000 matching recognition crops.
Each sample ID corresponds to:
one full image for text detection;
one cropped text image for text recognition;
one detector JSONL record;
one recognizer JSONL record.
Therefore, the dataset contains 240,000 aligned OCR pairs and
480,000 JPEG files in⦠See the full description on the dataset page:
https://huggingface.co/datasets/Phitran21/synthetic-ocr-en-det-rec-120k.