\ The ProofLang Corpus includes over three million
English-language proofs—558 million words—mechanically extracted from the papers
(Math, CS, Physics, etc.) posted on arXiv.org between 1992 and 2020. The focus
of this corpus is written proofs, not the explanatory text that surrounds them,
and more specifically on the language used in such proofs; mathematical
content is filtered out, resulting in sentences such as ``Let MATH be
the restriction of MATH to MATH.'' This dataset reflects how people prefer to
write informal proofs. It is also amenable to statistical analyses and to
experiments with Natural Language Processing (NLP) techniques.