This dataset contains Kannada document images with associated metadata including:
Translated layout information
Font information
VQA pairs
OmniDocBench annotations
image: PNG image files
file_name: Original filename
font_used: Font used in the document
page_size: Dimensions [width, height]
original_id: Original document identifier
rendered_layout: Layout elements with translations
translated_tables: Table translations… See the full description on the dataset page:
https://huggingface.co/datasets/v1v1d/vivid_augment_filtered.