The Thesis Dataset is a large-scale academic text corpus designed for building next-generation Natural Language Processing (NLP) systems, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) systems, research assistants, document understanding models, and knowledge extraction pipelines.
The complete multilingual collection contains over 641,700 thesis documents comprising 7.38+ billion words. This English release… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-Thesis-Dataset.