Openpdf-Blank-v2.0 is a small dataset containing blank or near-blank PDF image samples. This dataset is primarily designed to help train and evaluate document processing models, especially in tasks like:
Identifying and filtering blank or noise-filled documents.
Preprocessing stages for OCR pipelines.
Receipt/document classification tasks.
Modality: Image
Languages: English (if applicable)
Size: Less than 1,000 samples
License:… See the full description on the dataset page:
https://huggingface.co/datasets/prithivMLmods/Openpdf-Blank-v2.0.