D4LA is a large-scale, diverse, and manually annotated benchmark dataset for document layout analysis (DLA), introduced at ICCV 2023 alongside the Vision Grid Transformer (VGT) model. It is sourced from the real-world IIT-CDIP document collection and covers 12 document types with 27 fine-grained layout categories.
D4LA is considered the most diverse and richly annotated DLA benchmark, designed to reflect the complexity of real-world document… See the full description on the dataset page:
https://huggingface.co/datasets/tuandunghcmut/D4LA.