This dataset comprises authentic audio recordings collected from diverse locations within Hangzhou, China. It presents a valuable resource for researchers dedicated to advancing Speech-to-Text model performance. To ensure the privacy of individuals captured in the recordings, the dataset is encrypted. Researchers interested in accessing the dataset should contact the uploader, providing a detailed explanation of their research objectives and agreeing to strict confidentiality protocols.