Note: the dataset viewer is disabled — the data is distributed as zip bundles, not a
HuggingFace-loadable table. Download + unpack with the reproduce script (below), don't use load_dataset.
68,798 real image→tagged-text samples for teaching a VLM to emit inline formatting tags
(bold, italic, sup, sub, underline, strike). Sources: 22 arXiv categories (bold/italic/sup/sub) +
Washington & Texas legislative bills (underline/strike).… See the full description on the dataset page:
https://huggingface.co/datasets/ahamad-ai/mineru-decoder-finetune-dataset.