Benchmark of Tibetan Speech-To-Text dataset created by Monlam AI x Openpecha in 2024
All the transcripts have been reviewed by at least one person in addition to the original transcriber.
Data was taken on 15 July 2024 02∶47∶06 PM.
dept
desc
Count
STT_AB
Audio book
1000
STT_CS
Children Speech
1367
STT_HS
History
1000
STT_MV
Tibetan Movies
1000
STT_NS
Natural Speech
1000
STT_NW
News
1000
STT_PC
Podcast
1000
STT_TT
Tibetan Teachings
1000
grade column is used to… See the full description on the dataset page:
https://huggingface.co/datasets/openpecha/tibetan-voice-benchmark.