Pre-tokenized FineWeb data for language-model pretraining.
Source: HuggingFaceFW/fineweb
FineWeb configuration: sample-350BT
Tokenizer: tiktoken GPT-2 (gpt2)
Token dtype: little-endian uint16
Format: NanoGPT binary shards, version 1
Total tokens: 361,159,416,052
Shards: 1 validation shard and 3,611 training shards
Each .bin file contains a 1 KiB NanoGPT header followed by GPT-2 token… See the full description on the dataset page:
https://huggingface.co/datasets/tingtang2/fineweb350b-gpt2.