This dataset contains filtered images from the RVL-CDIP dataset, focusing on 4 specific document types.
A filtered subset of the RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset containing 100,000 images across 4 document categories. Each image is stored as base64-encoded data in Parquet format for efficient processing.
0
letter
Personal and… See the full description on the dataset page:
https://huggingface.co/datasets/sabaridsnfuji/rvl-cdip-filtered.