Views
No views yet
.oasr packs run with no Python at inference, engineered for peak performance on CPU & GPU1# 1. Install the OpenASR CLI · https://openasr.org
2# 2. Pull a build (pick a quant — see the table below)
3openasr pull whisper-large-v3:q8
4
5# 3. Transcribe
6openasr transcribe audio.wav --model whisper-large-v31openasr pull whisper-large-v3:fp16
2openasr pull whisper-large-v3:q8
3openasr pull whisper-large-v3:q4| Quant | File (.oasr) | Size | RAM peak | RTF · M1 CPU | RTF · M1 GPU | JFK ΔWER vs fp16 |
|---|---|---|---|---|---|---|
| fp16 | whisper-large-v3-fp16.oasr | 3.09 GB | 4.70 GB | 1.17× | 1.13× | 0.0% |
| q8_0 | whisper-large-v3-q8_0.oasr | 1.71 GB | 4.05 GB | 0.65× | 0.46× | 0.0% |
| q4_k | whisper-large-v3-q4_k.oasr | 1.29 GB | 2.78 GB | n/a | n/a | n/a |
openai/whisper-large-v3 weights as .oasr packs that run natively in the OpenASR runtime with
no Python at inference time. For most users the q8_0 build is the recommended default; q4_k is
for tighter memory budgets and fp16 is for verification or maximum fidelity. For a faster
large-grade option, see the distilled whisper-large-v3-turbo.1openasr model-pack import whisper <src> <out>.oasr \
2 --package-id whisper-large-v3 --quantization {fp16,q8-0,q4-k}.oasr container is GGUF-backed; packs use zero-copy mmap weight binding and graph
buffer reuse to keep peak memory low..oasr packages and adds quantized builds for local runtime use.