Dataset Card for eoinf/pile_gptneox
Original dataset
Original dataset: monology/pile-uncopyrighted
Dataset Details
Total Tokens: 200,897,536
Total Sequences: 196,189
Context Length: 1024 tokens
Tokenizer: EleutherAI/gpt-neox-20b
Format: Each example contains a single field tokens with a list of 1024 token IDs
Preprocessing
Each document was:
Tokenized using the EleutherAI/gpt-neox-20b tokenizer
Prefixed with a BOS (beginning of sequence) token… See the full description on the dataset page: https://huggingface.co/datasets/eoinf/pile_gptneox.