π FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from π FinePDFs dataset covering 69 languages.
FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset.
We then used this classifier to retain only the⦠See the full description on the dataset page:
https://huggingface.co/datasets/HuggingFaceFW/finepdfs-edu.