The training corpus for hofarah/orpheus-3b-persian-tts-lora:
19,458 Persian utterances, ≈50.8 hours of audio from a single female speaker, paired with
Finglish transcripts (Persian romanized into Latin script).
Each row is input_ids / labels / attention_mask — a fully assembled Orpheus training sequence.
There is no raw text column and no raw audio column. The audio has… See the full description on the dataset page:
https://huggingface.co/datasets/hofarah/Persian-tts-finglish-orpheus.