The Persian Document Corpus (PDC) is a large collection of Persian documents, comprising over 13,000 files, gathered from publicly accessible PDFs across a wide array of knowledge domains. This corpus includes research articles, theses, dissertations, scientific reports, and book chapters, offering a rich and diverse resource for the Persian Natural Language Processing (NLP) community. It is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-document-corpus.