Punjabi_ASR
Introduction
The
Punjabi_ASR project is dedicated to advancing Automatic Speech Recognition (ASR) for the Punjabi language, using various datasets to benchmark and improve performance. Our goal is to refine ASR technology to make it more accessible and efficient for speakers of Punjabi.
All Training, Evaluation, Processing scripts are available on
Github
Performance
We have benchmarked the ASR model using the IndicSuperb -
AI4Bharat/IndicSUPERB ASR benchmark with the following results:
- Common Voice: 10.18%
- Fleurs: 6.96%
- Kathbath: 8.30%
- Kathbath Noisy: 9.31%
These Word Error Rates (WERs) demonstrate the current capabilities and focus areas for improvement in our models.
Example Usage
To use the w2v-bert-punjabi model for speech recognition, follow the steps below. This example demonstrates loading the model and processing an audio file for speech-to-text conversion.
Code
1import speech_utils as su
2from m4t_processor_with_lm import M4TProcessorWithLM
3from transformers import Wav2Vec2BertForCTC, pipeline
4
5# Load the model and processor
6model_id = 'kdcyberdude/w2v-bert-punjabi'
7processor = M4TProcessorWithLM.from_pretrained(model_id)
8model = Wav2Vec2BertForCTC.from_pretrained(model_id)
9
10# Set up the pipeline
11pipe = pipeline('automatic-speech-recognition', model=model, tokenizer=processor.tokenizer, feature_extractor=processor.feature_extractor, decoder=processor.decoder, return_timestamps='word', device='cuda:0')
12
13# Process the audio file
14output = pipe("example.wav", chunk_length_s=20, stride_length_s=(4, 4))
15su.pbprint(output['text'])
Transcription:
ਉਹ ਕਹਿੰਦੇ ਸਾਡਾ ਸੁਨੇਹਾ ਹੁਣ ਜਾ ਕੇ ਅਹਿਮਦ ਸ਼ਾਹ ਬਦਾਲੀ ਨੂੰ ਦੇ ਦਿਓ ਉਹਨੇ ਸਾਨੂੰ ਪੇਸ਼ਕਸ਼ ਭੇਜੀ ਸੀ ਤਾਜ ਉਸ ਦਾ ਤੇ ਰਾਜ ਸਾਡਾ ਉਹਨੇ ਕਿਹਾ ਸੀ ਕਣਕ ਕੋਰਾ ਮੱਕੀ ਬਾਜਰਾ ਜਵਾਰ ਦੇ ਦਿਆ ਕਰੋ ਤੇ ਜ਼ਿੰਦਗੀ ਜੀ ਸਕਦੇ ਓ ਹੁਣ ਸਾਡਾ ਜਵਾਬ ਉਹਨੂੰ ਦੇ ਦਿਓ ਕਿ ਸਾਡੀ ਜੰਗ ਕੇਸ ਗੱਲ ਦੀ ਐ ਸਾਡੇ ਵੱਲੋਂ ਸ਼ਾਹ ਨੂੰ ਕਹਿ ਦੇਣਾ ਜਾ ਕੇ ਮਿਲਾਂਗੇ ਉਸ ਨੂੰ ਰਣ ਵਿੱਚ ਹੱਥ ਤੇਗ ਉਠਾ ਕੇ ਸ਼ਰਤਾਂ ਲਿਖਾਂਗੇ ਰੱਤ ਨਾਲ ਖੈਬਰ ਕੋਲ ਜਾ ਕੇ ਸ਼ਾਹ ਨਜ਼ਰਾਨੇ ਸਾਥੋਂ ਭਾਲਦਾ ਇਉਂ ਈਨ ਮਨਾ ਕੇ ਪਰ ਸ਼ੇਰ ਨਾ ਜਿਉਂਦੇ ਸੀਤਲਾ ਨੱਕ ਨੱਥ ਪਾ ਕੇ ਇਹ ਸੀ ਉਸ ਵੇਲੇ ਸਾਡੇ ਇਹਨਾਂ ਜਰਨੈਲਾਂ ਦਾ ਕਿਰਦਾਰ ਬਹੁਤ ਵੱਡਾ ਜੀਵਨ ਹੈ ਜਿਹਦੇ ਚ ਰਾਜਨੀਤੀ ਕੂਟਨੀਤੀ ਯੁੱਧ ਨੀਤੀ ਧਰਮਨੀਤੀ ਸਭ ਕੁਝ ਭਰਿਆ ਪਿਆ ਹੈ
Datasets
The training and testing data used in this project are available on Hugging Face:
Model
Our current model is hosted on Hugging Face, and you can explore its capabilities through the demo:
- Model: w2v-bert-punjabi
- Demo: Try the model
Next Steps
Here are the key areas we're focusing on to advance our Punjabi ASR project:
Collaboration and Support
We are actively seeking collaborators and sponsors to expand our efforts on the Punjabi ASR project. Contributions can be in the form of coding, dataset provision, or compute resources sponsorship. Your support will be crucial in making this practically beneficial for real-life applications.
- Issues and Contributions: Encounter an issue or want to help? Create a GitHub issue or submit a pull request to contribute directly.
- Sponsorship: If you are interested in sponsoring, especially in terms of compute resources, please email us at kdsingh.cyberdude@gmail.com to discuss collaboration opportunities.