Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
sea_speech – Dataset by gijs | AlphaNeural AI
You can deploy this model and start earning money today!
gijs
/
sea_speech
like
0
automatic-speech-recognition
audio-classification
zh
th
en
hi
vi
other
100K<n<1M
parquet
audio
text
datasets
dask
polars
mlcroissant
us
code-switching
speech
multilingual
Views
No views yet
Model card
Files and Versions
Community
API
SEA Code-Switching
101,975 verified code-switched speech clips · 473.74 hours · 42,321 distinct source videos, covering Chinese, Thai, English, Hindi and Vietnamese. Each clip contains at least one language switch by the same speaker, with the switch located in time.
By language pair
pair clips hours
en-hi 38,263 143.58
en-zh 37,964 207.57
en-th 15,035 75.45
en-vi 10,306 45.3
vi-zh 136 0.54
en-vi-zh 106 0.52
en-th-zh 73 0.39
th-zh 50 0.24… See the full description on the dataset page:
https://huggingface.co/datasets/gijs/sea_speech
.