This dataset will enable a object detection model to differentiate between shlok vs Non-Shlok in this types of books, after that shloks can be cropped and a parallel corpus of image-text can be created.
Dataset Details
Dataset Sources
PDF from internet archive was used, will link it here later