251,118 Turkish audio/text rows (~62 GB, 128 parquet shards). Schema is just
text + audio; see the YAML header above.
An early-phase Turkish speech collection from the TinyAya data pipeline. It is
not part of the v0.3 Stage-2 training corpus — that is
tr-hi-mimi-encoded.
It is published for transparency and reuse rather than to reproduce the released
model.
from datasets import load_dataset
ds = load_dataset("tiny-aya-translate/tr-subset-v0.1"… See the full description on the dataset page:
https://huggingface.co/datasets/tiny-aya-translate/tr-subset-v0.1.