ADOPD is a large-scale dataset designed for document image understanding. It introduces a novel data-driven document taxonomy discovery framework that combines large-scale pretrained models with a human-in-the-loop refinement process. The dataset supports four core tasks and includes rich annotations to foster progress in document analysis.
Total Images: 120,000
Languages:
English: 60… See the full description on the dataset page:
https://huggingface.co/datasets/adopd/adopd2024.