Synthetic Thai OCR dataset for fine-tuning vision-language models (e.g., DeepSeek-OCR-2).
Total samples: 4000
Subsets:
text — General Thai document text (1,000 samples)
table — Thai table formats: invoices, budgets, schedules, etc. (1,000 samples)
official — Thai government documents: contracts, legal, police reports, official letters (2,000 samples)
Parquet with embedded images. Compatible with HF dataset viewer.… See the full description on the dataset page:
https://huggingface.co/datasets/mekpro/ocr_th.