This dataset is a filtered, English-only subset of HuggingFaceFW/finepdfs-edu, created to retain high-signal educational passages while reducing common PDF-extraction noise (covers/TOCs, fragmented headers/footers, OCR artifacts, mixed-language pages, and very short low-context snippets).
It is intended for training and research workflows that benefit from longer, coherent educational text extracted from PDFs.
At a… See the full description on the dataset page: https://huggingface.co/datasets/C10X/finepdfs-edu-hq.