A large-scale, high-quality English text dataset derived from Project Gutenberg.
The corpus has been cleaned, normalized, deduplicated, segmented into fixed-length samples, and stored as compressed JSONL shards for efficient large-scale language model training.
This dataset is intended for pretraining and experimentation with small and medium language models such as TinyWay, tokenizer training, and large-scale NLP research.⦠See the full description on the dataset page:
https://huggingface.co/datasets/NNEngine/Gutenberg-Clean.