A deep learning automatic speech recognition (ASR) system for transcribing speech audio into IPA text using transformer-based sequence-to-sequence models.
1git clone https://github.com/neurlang/whipstr.git
2cd whipstr/
3uv run --with torch --with transformers --with phase-spectrogram stt_infer_hf.py --audio /home/m/Downloads/LJ001-0001.wav --model neurlang/ipa-whipstr-medium-48khz-cv-21
1config.json: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 368/368 [00:00<00:00, 2.09MB/s]
2model.safetensors: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 227M/227M [00:07<00:00, 29.4MB/s]
3Loading weights: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 199/199 [00:00<00:00, 16263.01it/s]
4model.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 3.84k/3.84k [00:00<00:00, 12.3MB/s]
5Transcription: ˈɪt gˈɪt sˈʌbkəldz dˈɑɹk fˈeɪz fˈæst ðə ɹˌɛpətˈɪʃən luːp nˈoʊ ˈiːvəz bɪˈheɪvjə ɹisˈɑlɪks lˈaɪk ðə mˈɑdəl hˈæd lˈɚnd bˈeɪsɪks ˈæktɪs....