This dataset contains tokenized speech for Orpheus fine-tuning.
Source dataset: dunkito/quijote
Number of examples: 65
Format: Each example contains input_ids, labels, and attention_mask in the format expected by Orpheus models.
This dataset is ready for fine-tuning Orpheus TTS models.
Created with the tokenise_speech_dataset.py script from the Trelis-Orpheus repository.