Public organic academic pretraining pool with 41,112,360 documents and
456,510,356,112 Dolma-2 tokens.
Common Pile peS2o Filtered multidisciplinary research papers
Dolma 3 olmOCR Science PDFs (score >=0.30, academic evidence, clean extraction)
Common Pile LibreTexts Filtered open textbook sections
quality_bin is null; source scores are retained only as source-specific evidence
labels: document type and publication year
exact dedup plus verified… See the full description on the dataset page:
https://huggingface.co/datasets/placeholderlabs/pool-academic.