Kleister Charity — Information Extraction
Dataset description and background
Kleister Charity is one of two corpora introduced in Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts (ICDAR 2021). It targets Key Information Extraction (KIE) from long, formally structured English documents that mix scanned pages and born-digital PDFs, with complex visual layout (tables, multi-column text, headers, etc.).
The Charity split is built… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/kleister_charity_information_extraction.