The LittleStories dataset is a collection of 5.4 million short stories generated by open-source language models. Inspired by roneneldan/TinyStories, this collection offers diverse narratives designed to teach text models about the world, relationships, and nuanced reasoning at a more realistic level.
The dataset is formatted in JSON for ease of use and split into manageable sizes for efficient processing.
Dataset Features… See the full description on the dataset page: https://huggingface.co/datasets/Corianas/LittleStories.