This dataset is a Japanese translation of the TinyStories dataset.
The dataset preserves the original TinyStories train/validation split structure and record order. Each record contains the original English story, its SHA-256 hash, and the Japanese translation.
text_hash: SHA-256 hash of text_en encoded as UTF-8.
text_en: Original English TinyStories text.
text_ja: Japanese translation.… See the full description on the dataset page:
https://huggingface.co/datasets/shibatch/TinyStories-JA.