RealText-V2 is a large-scale multilingual document benchmark dataset purpose-built for multilingual text image forgery analysis, pioneering in both scale and annotation depth.
Key Features
20K+ images: A large-scale benchmark, surpassing existing document forgery analysis datasets by orders of magnitude
6 languages: English, Chinese, Arabic, Thai, Malay, and… See the full description on the dataset page: https://huggingface.co/datasets/fang-123/RealText-V2.