Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
large-dataset-audio-v2 – Dataset by Tnaot | AlphaNeural AI
You can deploy this model and start earning money today!
Tnaot
/
large-dataset-audio-v2
like
0
automatic-speech-recognition
audio-classification
km
en
cc-by-nc-4.0
1K<n<10K
audiofolder
audio
datasets
mlcroissant
us
khmer
cambodian
speech
transcription
multilingual
Views
No views yet
Model card
Files and Versions
Community
API
khmer_speech_dataset
Khmer speech dataset with transcriptions, speaker labels, and metadata.
Dataset Description
This dataset contains Khmer (Cambodian) speech recordings with detailed transcriptions and annotations.
Dataset Statistics
Metric Value
Total Examples 9,285
Total Duration 336.68 hours
Average Duration 130.54 seconds
Total Words 1,336,611
Khmer Words 994,050 (74.4%)
English Words 340,127 (25.4%)
Unique Sources 2857
Unique… See the full description on the dataset page:
https://huggingface.co/datasets/Tnaot/large-dataset-audio-v2
.