Pre-tokenized version of HuggingFaceFW/fineweb-edu (sample-100BT config), prepared for autoregressive language model pretraining.
This is only the training split. Validation is wikitext-103 (not included here).
Tokenizer: gpt2
Sequence length: 1024
Wrap: True — documents are concatenated and chunked into fixed-length sequences
EOS inserted between documents
Saved with 🤗 datasets save_to_disk (Arrow… See the full description on the dataset page:
https://huggingface.co/datasets/mikezhu/fineweb-edu-100B-gpt2-tokenized-1024.