Views
No views yet
[!NOTE] 🌐 Live Web Demo: Try this model directly in your browser without any installation athuggingface.co/spaces/rudrakshrakeshzodage/neo-whisper
[!NOTE] This model is part of a suite of optimized/quantized versions of the base model. Other variants in this direction:
- PyTorch Q4 (GPU Quantized - NF4):
rudrakshrakeshzodage/whisper-large-v3-turbo-pytorch-q4(Current)- CTranslate2 INT8 (CPU Quantized):
rudrakshrakeshzodage/whisper-large-v3-turbo-ct2-int8- Browser ONNX Q4 (WebGPU/Wasm):
rudrakshrakeshzodage/whisper-large-v3-turbo-onnx
openai/whisper-large-v3-turbo

torch.nn.Linear layers in Whisper self-attention, cross-attention, and MLP blocks.bitsandbytes to perform dynamic matrix multiplication.| Quantization Format | Target Backend | Target Device | Average WER (%) |
|---|---|---|---|
| CTranslate2 Float32 (CPU) | CTranslate2 | CPU | 45.62% |
| CTranslate2 INT16 (CPU) | CTranslate2 | CPU | 45.96% |
| CTranslate2 INT8 (CPU) | CTranslate2 | CPU | 46.01% |
| PyTorch FP16 (GPU) | PyTorch | GPU | 44.73% |
| PyTorch INT8 (GPU) | PyTorch | GPU | 41.20% |
| PyTorch Q4 (GPU) | PyTorch | GPU | 44.50% |
| Language | PyTorch FP16 (GPU) | PyTorch INT8 (GPU) | PyTorch Q4 (GPU) | CTranslate2 Float32 (CPU) | CTranslate2 INT8 (CPU) | CTranslate2 INT16 (CPU) |
|---|---|---|---|---|---|---|
| English | 63.0% | 63.0% | 63.0% | 69.6% | 69.6% | 69.6% |
| Spanish | 41.1% | 41.1% | 41.1% | 41.1% | 41.1% | 41.1% |
| French | 45.8% | 45.8% | 47.3% | 47.3% | 47.3% | 47.3% |
| Italian | 29.9% | 32.5% | 35.0% | 29.9% | 29.9% | 32.5% |
| German | 18.8% | 18.8% | 22.1% | 18.8% | 18.8% | 18.8% |
| Chinese | 100.0% | 66.7% | 100.0% | 100.0% | 100.0% | 100.0% |
| Korean | 38.2% | 43.0% | 38.2% | 40.8% | 41.8% | 40.8% |
| Dutch | 33.3% | 33.3% | 33.3% | 33.3% | 31.2% | 33.3% |
| Russian | 51.4% | 44.7% | 47.8% | 48.6% | 55.3% | 48.6% |
| Czech | 25.8% | 23.3% | 17.2% | 26.8% | 25.0% | 27.6% |
1import torch
2from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, BitsAndBytesConfig
3
4model_id = "rudrakshrakeshzodage/whisper-large-v3-turbo-pytorch-q4"
5
6quant_config = BitsAndBytesConfig(
7 load_in_4bit=True,
8 bnb_4bit_quant_type="nf4",
9 bnb_4bit_use_double_quant=True,
10 bnb_4bit_compute_dtype=torch.float16
11)
12
13model = AutoModelForSpeechSeq2Seq.from_pretrained(
14 model_id,
15 quantization_config=quant_config,
16 device_map="auto"
17)
18processor = AutoProcessor.from_pretrained(model_id)