AncientDoc is the first comprehensive benchmark dataset specifically designed for Chinese Ancient Document Understanding. It covers multi-task evaluation ranging from OCR to knowledge reasoning, aiming to promote research on the recognition, understanding, and reasoning capabilities of multimodal large models in the scenario of ancient documents.
Data Scale: 2,973 pages
Number of… See the full description on the dataset page:
https://huggingface.co/datasets/ByteDance/AncientDoc.