Pre-extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities.
This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT-2 tokens) extracted from the raw source corpus (PDF, EPUB, TEI XML, HTML). If you need the original source formats, see the raw corpus.