Project Gutenberg is an online collection of over 75,000 digitized books available as plain text.
We use all books that are 1) English and 2) marked as in the Public Domain according to the provided metadata.
Additionally, we include any books that are part of the PG19 dataset, which only includes books that are over 100 years old.
Minimal preprocessing is applied to remove the Project Gutenberg header and footers, but many scanned books… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/project_gutenberg_filtered.