DocBank is a new large-scale dataset that is constructed using a weak supervision approach. It enables models to integrate both the textual and layout information for downstream tasks. The current DocBank dataset totally includes 500K document pages, where 400K for training, 50K for validation and 50K for testing.
[GitHub] [Paper]
We update the license to Apache-2.0.
Our paper has been accepted in COLING2020 and the Camera-ready version paper has been… See the full description on the dataset page:
https://huggingface.co/datasets/liminghao1630/DocBank.