This is a subset of the RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset, containing 200 samples per class for a total of 3,200 samples. The dataset consists of scanned document images in TIFF format, collected from various sources. The documents belong to 16 different categories, such as letter, memo, email, and more. The purpose of this dataset is to facilitate document classification tasks using… See the full description on the dataset page: https://huggingface.co/datasets/hadasBublil/rvl_cdip-small-200.