This dataset is a highly compressed, pre-tokenized subset designed exclusively for educational purposes and university-level AI coursework. It provides a lightweight sandbox for students to explore compute-optimal scaling, tokenizer compression, and language model training in heavily constrained environments.
To maximize batching efficiency, all stories have been strictly filtered to a Maximum Sequence… See the full description on the dataset page:
https://huggingface.co/datasets/ChemCogLab/tinier_stories.