This is a preprocessed, tokenized dataset for the cramming-project.
Use only with the tokenizer uploaded here.
This version is 97b8e776baafb99c3892e6572a9f51b3, which corresponds to a specific dataset construction setup, described below.
The raw data source is the Pile, a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality
datasets combined together.… See the full description on the dataset page:
https://huggingface.co/datasets/JonasGeiping/the_pile_WordPiecex32768_97b8e776baafb99c3892e6572a9f51b3.