This dataset contains preprocessed and token-packed .bin files intended for use in pretraining a decoder-only Transformer language model.
Each .bin file contains a fixed number of samples, where each sample is exactly 8192 tokens long.
Samples are grouped into batches of 125 samples, totaling 1.024 million tokens per batch.
Each file (called a "block") contains 62500 samples (approximately 512 million tokens).
All… See the full description on the dataset page:
https://huggingface.co/datasets/wottAI/textpack-20b-tokenized.