math-reasoning-packed-1024
Packed token dataset ready for GPT-style pretraining.
This dataset contains pre-tokenized and packed sequences optimized for efficient transformer training.
Encoding: cl100k_base
Vocabulary Size: 100,277
Sequence Length: 1024
Documents: 684,238
Total Tokens: 239,019,783
Sequences: 233,417
Shards: 3
Documents: 13,961
Total Tokens: 4,875… See the full description on the dataset page:
https://huggingface.co/datasets/ethanker/math-reasoning-packed-1024.