Project-Gutenberg-enPurified is a highly curated, "prose-first" refinement of the Project Gutenberg corpus.
The enPurified collection is built on a specific philosophy: Specialization. While most modern datasets are "general purpose," they often dilute linguistic quality with code snippets, math formulas, and broken OCR text. This dataset aggressively strips away everything but high-quality English prose to help models master fluid… See the full description on the dataset page:
https://huggingface.co/datasets/enPurified/project_gutenberg-enPurified-openai-messages.