Open-source subset of the Chunkr Layout Detection benchmark dataset, containing 360 professionally annotated documents with 4,695 layout annotations across 16 element classes.
This dataset benchmarks document layout detection models on diverse, real-world documents including financial reports, legal contracts, research papers, medical records, and more.It draws from a curated collection of publicly accessible or research datasets listed below… See the full description on the dataset page:
https://huggingface.co/datasets/ChunkrAI/chunkr-layout-bench-oss.