Ultra-FineWeb-en 20B GPT-4 Token Shards
Binary token shards for LR-AttnRes.
Source dataset: openbmb/Ultra-FineWeb, split en
Text column: content
Tokenizer model: gpt-4
Encoding: cl100k_base
Vocab size: 100277
Document separator / EOT token: 100257
Dtype: uint32
The shard filenames intentionally use the existing dataloader convention:
finewebedu_val_.bin
finewebedu_train_.bin
Download with:
python prepdata.py --repo-id /Ultra-FineWeb-en-20B-gpt4