Cleaned, deduplicated, and split version of roneneldan/TinyStories,
prepared for training AnkurLM (a from-scratch GPT-style small language model).
This dataset is a modified/derived version of TinyStories (roneneldan/TinyStories),
which is licensed under CDLA-Sharing-1.0. Per that license's share-alike terms,
this derived dataset is also released under CDLA-Sharing-1.0.
Unicode normalization… See the full description on the dataset page:
https://huggingface.co/datasets/Subhadip007/AnkurLM-Dataset.