This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, with tokenized book texts using the tiktoken tokenizer gpt2 encoding.
Total number of tokens is 2,110,010.
It was created for learning, testing streaming datasets, and quick downloading and manipulation.
It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset.
The text of the books… See the full description on the dataset page:
https://huggingface.co/datasets/nikolina-p/mini_gutenberg_tokenized.