Tagin Speech Corpus (Taliha Circle)
Dataset Description
The Tagin Speech Corpus is a raw, uncurated collection of speech recordings in the Tagin language (ISO 639-3: tgj). The data was primarily sourced from native speakers conversing in WhatsApp groups belonging to the Taliha Circle region in the Upper Subansiri district of Arunachal Pradesh, India.
This dataset is intended to support the development and research of Automatic Speech Recognition (ASR) systems… See the full description on the dataset page: https://huggingface.co/datasets/repleeka/Tagin-Speech-Corpus.