š Paper | š Project Page | š¾ Released Resources | š¦ Repo
We release the raw transcription data for our SynthVoice project, adapted from the original LibriSpeech dataset by OpenSLR.
The data format for each line in the voice_transcripts_raw.jsonl is as follows:
{
"audio_id": "
","transcript": "",
"speaker_id": "",
"duration_seconds":⦠See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/SynthVoice-Raw.