Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
omniasr-molge – Dataset by Sanghyang00 | AlphaNeural AI
You can deploy this model and start earning money today!
Sanghyang00
/
omniasr-molge
like
0
automatic-speech-recognition
other
1M<n<10M
parquet
tabular
text
datasets
dask
polars
mlcroissant
2511.09690
2606.11681
2607.24030
us
Views
No views yet
Model card
Files and Versions
Community
API
OmniASR Molge Aligned
Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC.
Source facebook/omnilingual-asr-corpus
Configs omniasr_aligned_v1, omniasr_aligned_v2
Splits train / validation (dev-*.parquet) / test
Scale ~2.56M utts · ~839 shards · ~439GB
If this dataset is useful for your work, we’d appreciate a… See the full description on the dataset page:
https://huggingface.co/datasets/Sanghyang00/omniasr-molge
.