The ATR benchmark dataset is a multilingual dataset that includes 463 document images, at page or paragraph level. This dataset has been designed to test ATR models and combines data from several public datasets:
BnL Historical Newspapers
CASIA-HWDB2
Churro
DIY History - Social Justice
DAI-CRETDHI
Esposalles
FINLAM - Historical Newspapers
Horae - Books of hours
IAM
NDLOCR
NorHand v3
OpenITI
Marius PELLET
QARI
RASM… See the full description on the dataset page: https://huggingface.co/datasets/Teklia/ATR-benchmark.