Same format as authentic dataset here
Added train split (70%) and test split (30%)
authentic data of the same split could be found on authentic dataset with split
synthesized using Openai tts (self-funded)
based on transcription of the anthentic dataset
synthesized using TTS for mandarin, with no external prompt
For non-comercial use only