Openpdf-Blank-v2.0-Sample is a sample dataset of blank or near-blank invoice and receipt documents. It contains 255 high-resolution scanned images extracted and cleaned from document PDFs. This dataset is intended to support training and evaluation of OCR, document classification, and layout-based filtering models where blank or structurally minimal pages must be identified and processed.
Format: Parquet (auto-converted)… See the full description on the dataset page:
https://huggingface.co/datasets/prithivMLmods/Openpdf-Blank-v2.0-Sample.