PDF documents are typically paginated, which often results in tables or paragraphs being split across consecutive pages. Accurately detecting and merging such cross-page structures is crucial to avoid generating incomplete or fragmented content.
The merging of two table fragments is especially challenging. For example, the table spanning multiple pages will repeat the header of the first page on the second page. Another difficult scenario is that the table… See the full description on the dataset page:
https://huggingface.co/datasets/ChatDOC/OCRFlux-pubtabnet-cross.