This dataset contains processed research papers optimized for GPT-NeoX-20B training.
The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling.
Total Samples: 9993
Unique Papers: 1017
Average Tokens per Sample: 1965.4
Token Range: 10 - 91659
Max Token Limit: 2048
Source Subdirectories: 1
text: The processed research paper text or chunk… See the full description on the dataset page:
https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.