A high-quality, pre-tokenized, length-bucketed derivative of
HuggingFaceFW/finepdfs-edu,
prepared as a drop-in corpus for distributed continuous / continual pretraining (CPT).
The source documents are filtered to confident English, tokenized once with the
jet-ai/Jet-Nemotron-2B tokenizer, and written to
flat uint32 token shards that are partitioned by sequence length so that training jobs can pull
short or long sequences on… See the full description on the dataset page:
https://huggingface.co/datasets/tturing/finepdfs-edu-32.