This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research.
Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data.
The dataset consists of a single split:… See the full description on the dataset page:
https://huggingface.co/datasets/Nelathan/standardebooks.