Search 3.2M models and datasets…
⌘K
Chat
Models
Datasets
Deploy
Pricing
Docs
Chat
Models
Datasets
Deploy
More
zoengjyutgaai – Dataset by CanCLID | AlphaNeural AI
Is this your dataset? Claim it with the Hugging Face account that owns it.
CanCLID
/
zoengjyutgaai
like
0
automatic-speech-recognition
text-to-speech
text-generation
feature-extraction
audio-to-audio
audio-classification
text-to-audio
yue
cc0-1.0
100K<n<1M
parquet
audio
text
datasets
dask
polars
mlcroissant
us
cantonese
audio
Views
No views yet
Dataset card
Files and Versions
Community
Use
Use this dataset
張悦楷講古語音數據集
English 呢個係張悦楷講《三國演義》、《水滸傳》、《走進毛澤東的最後歲月》、《鹿鼎記》語音數據集。張悦楷係廣州最出名嘅講古佬 / 粵語説書藝人。佢從上世紀七十年代開始就喺廣東各個收音電台度講古,佢把聲係好多廣州人嘅共同回憶。本數據集收集嘅係佢最知名嘅四部作品。 數據集用途:
TTS(語音合成)訓練集 ASR(語音識別)訓練集或測試集 各種語言學、文學研究 直接聽嚟欣賞藝術!
TTS 效果演示:
https://huggingface.co/spaces/laubonghaudoi/zoengjyutgaai_tts
説明
所有文本都根據
https://jyutping.org/blog/typo/
同
https://jyutping.org/blog/particles/
規範用字。 所有文本都使用全角標點,冇半角標點。 所有文本都用漢字轉寫,無阿拉伯數字無英文字母 所有音頻源都存放喺/source,為方便直接用作訓練數據,切分後嘅音頻都放喺 opus/ 所有 opus 音頻皆為 48000… See the full description on the dataset page:
https://huggingface.co/datasets/CanCLID/zoengjyutgaai
.