DIT700K is a new benchmark dataset for document image translation built with large-scale data examples and fine-grained multi-level labels, enabling the research on En-Zh/De and Zh-En DIT, reading order detection, and layout analysis. Note that the DITrans dataset can only be used for non-commercial research purposes. For scholars or organizations who want to use the dataset, please send an application via email to us (
zhangzhiyang2020@ia.ac.cn). When submitting the… See the full description on the dataset page:
https://huggingface.co/datasets/zhangzhiyang/DIT700K.