This dataset contains text extracted from a collection of 160 Nigerian authored books.
Curated by: Abdullahi Mujaheed, Ayeni Oluwatosin, Nworie Kingsley
Language(s) (NLP): en
License: apache-2.0
The dataset consists of a single "train" split. Each example in the dataset contains a single feature:
text: A string value representing a single line of raw text extracted from the PDF documents. Each line… See the full description on the dataset page:
https://huggingface.co/datasets/MLHermit/9ja-bookcorpus.