This dataset contains:
40351 images (71.39%) in train dataset
6378 images (14.29%) in validation dataset
6391 images (14.32%) in test dataset
Total: 53120 images
sourced from the extensive Synthetic Word Dataset, a large-scale word-image dataset.
The original and complete dataset (9 million images, 10.68GB) can be found and downloaded at this academic torrent.