In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively.
Please refer to our blogpost and paper (Coming soon!) for further details.
To load and use dataset, run this script:… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.