OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged
for corpus exploration, language modeling, OCR-quality analysis, and historical
Ottoman Turkish / Turkish NLP work.
pages: one row per OCR page, with page numbers and OCR status.
documents: one row per available volume, with page text concatenated.
Available books: 1, 3, 4, 6, 7, 9, 10.
Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page:
https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.