Openpdf-MultiReceipt-1K is a dataset consisting of over 1,000 receipt documents in PDF format. This dataset is designed for use in image-to-text and document understanding tasks, particularly Optical Character Recognition (OCR), receipt parsing, and layout analysis.
No text annotations or metadata are provided — only the raw PDFs.
Ideal for tasks requiring raw document inputs like PDF-to-Text pipelines.