This Whisper-for-IPA (WhIPA) model is a fine-tuned version of openai/whisper-base on a subset of the CommonVoice11 dataset (1k samples each from Greek, Finnish, Hungarian, Japanese, Maltese, Polish, Tamil) with G2P-based IPA transcriptions.
It (Ckpt4) achieves the following results on the evaluation set: