CamStories-Full is the full cleaned and combined corpus of children’s fiction stories from the TinyStories GPT-4 variant and SimpleStories datasets.
It uses the same extensive cleaning, name normalisation, and grammar filtering pipeline as CamStories-10k, but retains the full uncased word vocabulary of 31,244 tokens. This makes it a superset of CamStories-10k, preserving more lexical variety while still offering consistent formatting and per-token audio.… See the full description on the dataset page:
https://huggingface.co/datasets/Piros/CamStories_full.