I have created a dataset containing 50 short Punjabi sentences, totaling about 3.5 minutes of speech. I recorded all the sentences myself in a quiet room using a consistent microphone setup to ensure clear and natural audio. Each sentence has been exported as a separate WAV file and is aligned with its transcription in metadata.csv. The dataset is designed to be clean and well-segmented, making it suitable for Text-to-Speech (TTS) and other speech processing… See the full description on the dataset page:
https://huggingface.co/datasets/Sukhi2144/11500512_KaurSukhvir.