Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
uyghur-whisper-finetune – Dataset by anke01 | AlphaNeural AI
You can deploy this model and start earning money today!
anke01
/
uyghur-whisper-finetune
like
0
automatic-speech-recognition
translation
ug
zh
cc-by-4.0
1K<n<10K
audiofolder
audio
datasets
mlcroissant
us
uyghur
chinese
speech
whisper
low-resource
parallel-corpus
Views
No views yet
Model card
Files and Versions
Community
API
UyZh-FolkSpeech Whisper 微调数据集
维吾尔语-汉语平行语音数据集,适用于 Whisper 模型微调。
数据集来源
本数据集源自 UyZh-FolkSpeech,经过以下处理:
音频格式转换: M4A → WAV (16kHz, 单声道) 文本清洗: 移除不可见控制字符 (U+200E, U+200F 等) 数据集划分: train/val/test (80%/10%/10%) 目录重组: 按划分分文件夹存储
数据统计
划分 记录数 音频数 时长
train 6,348 6,348 407.60 分钟
validation 793 793 48.89 分钟
test 795 795 51.80 分钟
总计 7,936 7,936 508.29 分钟
内容分布
短句 (sentence): 3,812 条 词汇短语 (word): 4,124 条
说话人分布
每个文本条目由 4 位说话人… See the full description on the dataset page:
https://huggingface.co/datasets/anke01/uyghur-whisper-finetune
.