The Harvard USPTO Patent Dataset (HUPD) is a large-scale, well-structured, and multi-purpose corpus
of English-language patent applications filed to the United States Patent and Trademark Office (USPTO)
between 2004 and 2018. With more than 4.5 million patent documents, HUPD is two to three times larger
than comparable corpora. Unlike other NLP patent datasets, HUPD contains the inventor-submitted versions
of patent applications, not the final versions of granted patents, allowing us to study patentability at
the time of filing using NLP methods for the first time.