This dataset contains text extracted from pharmaceutical PDF sources,
used for domain-adaptive continued pretraining of a causal language
model (TinyLlama).
Files
pdf_pages_raw.jsonl — raw text extracted page-by-page from source PDF(s),
before any cleaning.
pharma_paragraph_process.jsonl — cleaned and paragraph-split version of
the raw text, used directly for tokenization and training.