MinSpeech: Cleaned Multi-dialect Min-nan Dataset (Private)
Important Legal Notice & Copyright Status
This repository is a Private Research Fork of the MinSpeech corpus. It is maintained strictly for individual research purposes, specifically for fine-tuning Automatic Speech Recognition (ASR) and Speech-to-Text Translation (S2TT) models.
1. Ownership & Licensing
Annotations & Metadata: The transcriptions and segment metadata are derived from the MinSpeech… See the full description on the dataset page: https://huggingface.co/datasets/scbz/minspeech.