Dataset Card for Stanza-TinyStories-2
Dataset Summary
Stanza-TinyStories-2 is a structurally and morphologically enriched iteration of the TinyStories dataset (Eldan and Li, 2023).
This dataset projects the 1D synthetic text generated by large language models into a fully resolved grammatical and topological space. Every sentence in the 2.7-million-story training split and the 21,000-story validation split has been deterministically parsed to extract Universal… See the full description on the dataset page: https://huggingface.co/datasets/EXOROBOURII/Stanza-TinyStories.