If you are looking for a program to run this model with, then I would recommend
EasyWhisper UI, as it is user-friendly, has a GUI, and will automate a lot of the hard stuff for you.
Clicking on a link will download the corresponding quant instantly.
My guess is that your GPU might be too old to recognize them, considering that I have gotten the same error on my GTX 1080. If you would like to run them regardless, you can try switching to CPU inference.
The quantizer I was using was not specific about this, so I do not know about this either.
I used
whisper.cpp v1.7.6 on Windows x64, leveraging CUDA 12.4.0. For the F32 quant, I converted the original Hugging Face (H5) format model to a GGML using the
models/convert-h5-to-ggml.py script.
Open a new discussion in the community tab about this, and I will look into the issue.