For more information, please visit
https://github.com/clovaai/donut
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
If you find this work useful… See the full description on the dataset page:
https://huggingface.co/datasets/naver-clova-ix/synthdog-ko.