A VIVID dataset containing document layout information with rich annotations. :)
This dataset contains PDF document pages with extracted layout information including:
Layout boxes: Bounding boxes for text, images, tables, equations
Multiple output formats: Markdown, HTML, HTML with coordinates
VQA pairs: Visual question-answer pairs for each page
Raw content: Lines, images, equations, tables extracted separately… See the full description on the dataset page:
https://huggingface.co/datasets/v1v1d1/vivid_docmatix_final_250k.