TDoc-2.8M is a large-scale multilingual dataset for tampered text detection and localization in document images.
It contains approximately 2.8 million tampered document images from multiple languages. The dominant languages are English, French, and Chinese. The full dataset size is about 7.44 TB.
The dataset is distributed as compressed shards:
shards/
shard_000000.tar.zst
shard_000001.tar.zst
...
Download… See the full description on the dataset page:
https://huggingface.co/datasets/MohamedDhouib1/TDoc-2.8M.