Audar-ASR-V1-Turbo is an Arabic-first generative speech-recognition model — the accuracy tier of
the Audar-ASR family. It recasts transcription as audio-conditioned next-token prediction over a
unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is
built on a permissively-licensed open-weight audio-LLM foundation and adapted in-house — the contribution is the adaptation (the data curriculum and the alignment rubric), not the foundation:
🧱 Large-scale bilingual pretraining — 300,000+ hours of labeled audio, primarily Arabic and English, spanning MSA, Gulf,
Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
🎯 Dialect-targeted fine-tuning — hardness sampling and multi-task conditioning focused on proper
nouns, code-switching, and dialect-faithful orthography.
🧠 KTO preference alignment — Kahneman-Tversky Optimization on accented dialectal Arabic, with
unpaired binary-desirability labels from trained native annotators across the Gulf, Levantine,
Egyptian, and Maghrebi dialects, along five axes: verbatim accuracy, diacritic correctness,
code-switch handling, named-entity preservation, and output formatting.
The result is state-of-the-art dialectal Arabic ASR — the lowest average WER and CER of any
evaluated system on the Open Universal Arabic ASR Leaderboard. It transcribes MSA and every major
Arabic dialect, code-switched Arabic–English, and English, across 30 languages in total.
Built on a permissively-licensed open-weight audio-LLM foundation; the adaptation, data, and
alignment are Audar's. Full method and results: Audar-ASR-V1 Technical Report.
Model summary
Model
Audar-ASR-V1-Turbo — Arabic-first generative ASR (accuracy tier)
built on an open-weight audio-LLM foundation; adapted via a 4-stage curriculum — 300k+ hrs bilingual pretraining → multi-task fine-tuning → dialect PEFT → KTO alignment
Decoder parameters
2,031,739,904 (2.03B)
Audio encoder parameters
317,477,504 (0.32B)
Total parameters
2,349,217,408 (2.35B, bf16)
Audio input
16 kHz mono; 30 s context (longer audio is chunked/streamed)
Languages
Arabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more
Runtime
GGUF / llama.cpp — CPU · GPU · edge
License
AudarAI Community License v1.0
📊 Benchmarks
Arabic dialectal ASR is hard — heavily dialectal, conversational, code-switched speech is the
frontier for every system. On the Open Universal Arabic ASR Leaderboard, Audar-ASR-V1-Turbo ranks
#1 of 37 systems with the lowest average WER (23.2 %) and the lowest average CER (9.2 %) of any
model evaluated — and it is the single best system on
SADA, MASC-clean, MGB-2 and Casablanca.
Open Universal Arabic ASR Leaderboard — full standings
Per-dataset WER % across all six leaderboard test sets, plus the two composite averages. Lower is
better; Avg WER is the ranking metric. Audar rows show the leaderboard maintainers' independent reproduction (Aug 2026) under the
leaderboard's current normalization; other rows are as previously published by the leaderboard and
may shift slightly when the full board is recomputed under the updated normalization. Ours in bold.
#
Model
Avg WER
Avg CER
SADA
CV-18
MASC-clean
MASC-noisy
MGB-2
Casablanca
1
Audar-ASR-V1-Turbo (Ours)
23.17
9.20
28.92
8.09
16.73
27.19
11.08
47.02
2
CohereLabs/cohere-transcribe-arabic-07-2026
25.87
11.80
37.47
5.82
19.60
27.07
15.54
49.71
3
omnilingual-asr/omniASR_LLM_7B
28.32
12.52
41.61
8.75
19.69
29.29
14.13
56.46
4
omnilingual-asr/omniASR_LLM_3B
29.96
13.77
46.18
9.15
19.90
30.03
14.22
60.27
5
omnilingual-asr/omniASR_LLM_1B
29.96
13.40
43.84
9.55
20.03
30.26
15.34
60.68
6
CohereLabs/cohere-transcribe-03-2026
30.67
16.37
60.11
8.17
8.66
19.01
25.33
62.71
7
Qwen/Qwen3-Omni-30B-A3B-Instruct
30.71
13.67
44.82
11.46
21.47
30.85
13.09
62.55
8
nvidia-conformer-ctc-large-arabic (lm)
32.91
13.84
44.52
8.80
23.74
34.29
17.20
68.90
9
omnilingual-asr/omniASR_LLM_300M
32.96
14.84
51.38
12.03
20.66
32.45
16.58
64.64
10
google/gemma-4-E4B-it
32.98
13.71
43.40
19.65
24.86
33.59
17.72
58.63
11
Qwen/Qwen3-ASR-1.7B
33.36
12.33
45.53
16.90
24.37
34.29
16.57
64.47
12
mistralai/Voxtral-Small-24B-2507
34.47
15.29
50.82
15.25
23.96
34.43
16.03
66.30
13
nvidia-conformer-ctc-large-arabic (greedy)
34.74
13.37
47.26
10.60
24.12
35.64
19.69
71.13
14
google/gemma-4-E2B-it
35.87
15.34
46.23
23.76
27.47
36.15
20.72
60.87
15
openai/whisper-large-v3
36.86
17.21
55.96
17.83
24.66
34.63
16.26
71.81
16
omnilingual-asr/omniASR_CTC_3B
37.78
19.79
69.85
14.19
21.48
34.60
18.96
67.58
17
omnilingual-asr/omniASR_CTC_7B
38.12
20.91
72.69
12.47
21.08
35.04
20.43
67.02
18
facebook/seamless-m4t-v2-large
38.16
17.03
62.52
21.70
25.04
33.24
20.23
66.25
19
omnilingual-asr/omniASR_CTC_1B
39.29
20.47
71.42
17.55
22.76
35.73
19.96
68.32
20
openai/whisper-large-v3-turbo
40.05
18.87
60.36
25.73
25.51
37.16
17.75
73.79
21
openai/whisper-large-v2
40.20
19.55
57.46
21.77
27.25
38.55
25.17
71.01
22
Qwen/Qwen3-ASR-0.6B
42.19
16.23
53.75
28.28
31.34
42.63
25.45
71.68
23
openai/whisper-large
42.57
20.49
63.24
26.04
28.89
40.79
24.28
72.18
24
mistralai/Voxtral-Mini-3B-2507
42.58
19.90
63.65
22.12
28.37
41.27
22.56
77.52
25
asafaya/hubert-large-arabic-transcribe
45.50
17.35
67.82
8.01
32.94
50.16
37.51
76.53
26
openai/whisper-medium
45.57
22.27
67.71
28.07
29.99
42.91
29.32
75.44
27
nvidia-Parakeet-ctc-1.1b-concat
46.54
23.88
70.70
26.34
30.49
45.95
24.94
80.80
28
omnilingual-asr/omniASR_CTC_300M
46.65
21.86
78.11
27.90
28.40
43.26
26.85
75.35
29
nvidia-Parakeet-ctc-1.1b-universal
51.96
25.19
73.58
40.01
36.16
50.03
30.68
81.30
30
microsoft/VibeVoice-ASR
52.99
28.95
69.83
44.25
32.95
52.43
25.10
93.37
31
facebook/mms-1b-all
54.54
21.45
77.48
26.52
38.82
57.33
39.16
87.95
32
openai/whisper-small
55.13
21.68
78.02
24.18
35.93
56.36
48.64
87.64
33
whitefox123/w2v-bert-2.0-arabic-4
58.13
27.62
87.34
41.79
37.82
53.28
40.66
87.88
34
jonatasgrosman/wav2vec2-large-xlsr-53-arabic
60.98
25.61
86.82
23.00
42.75
64.27
56.29
92.72
35
speechbrain/asr-wav2vec2-commonvoice-14-ar
65.74
30.93
88.54
29.17
49.10
69.57
64.37
93.68
Bold = best in column. The 37th system, our sibling edge model Audar-ASR-V1-Flash (0.78B), enters at 32.04 avg WER — see its card for the full row. Audar-ASR-V1-Turbo owns both composite averages and leads on SADA, MASC-clean, MGB-2 and Casablanca;
the recent Cohere and OmniASR systems are the closest competitors, each strongest on a subset of the
conversational and clean-read sets. Casablanca (Moroccan Darija) is the hardest set for every system.
Emirati Arabic
Set
WER %
CER %
Emirati (Mixat, full 1,585-clip test)
19.4
7.3
On Emirati, the real recognition error is ≈ 7.3 % — near-parity with spontaneous English — while the
residual up to 19.4 % WER is largely orthographic convention (near-miss spelling of the same
word, e.g. انتو↔انتوا, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.
Our leaderboard numbers were produced with the qwen-asr package, which implements this model's I/O protocol natively — and were independently reproduced by the leaderboard maintainers with this exact code:
language="Arabic" makes the package prefill language Arabic<asr_text> into the prompt, so the model never free-runs language identification.
The model's no-speech verdict (language None<asr_text>) is mapped to an empty transcript; without this, non-speech audio (music, silence) can produce repetition loops.
max_new_tokens=256 and bf16 are the exact decode settings behind our published numbers.
If you use raw transformers (below), you must strip the language <Lang><asr_text> output prefix yourself and expect degraded scores on non-speech-heavy data.
💻 GGUF inference (llama.cpp)
Turbo runs on llama.cpp via the multimodal (mtmd) path — a quantized decoder GGUF plus a
BF16 audio projector (mmproj). Build a recent llama.cpp (with Qwen3-ASR support), then:
⚠️ The audio projector (mmproj) must stay BF16 (its ClippableLinear is numerically
sensitive). The decoder quantizes normally.
Prefer a managed endpoint? The Audar-ASR family is also available via the
Audar API/SDK — streaming, speaker-attributed transcription, and
diarization, production-hosted.
GGUF variants
File
Approx. size
Notes
Audar-ASR-V1-Turbo-Q4_K_M.gguf
~1.28 GB
Smallest; constrained hardware
Audar-ASR-V1-Turbo-Q8_0.gguf
~2.16 GB
Near-lossless (recommended)
Audar-ASR-V1-Turbo.gguf (BF16)
~4.07 GB
Full precision decoder
mmproj-Audar-ASR-V1-Turbo.gguf
~0.64 GB
BF16 audio encoder — required, keep BF16
🤗 Transformers (full-precision safetensors)
The full-precision bf16 weights are published at the repo root — the reference
checkpoint the GGUF and W4A16 builds are derived from (2,349,217,408 params, safetensors). Standard
🤗 Transformers, loaded with trust_remote_code=True (the repo ships the self-contained Qwen3-ASR code).
python
1# pip install "transformers==4.57.6" torch librosa2import torch, librosa
3from transformers import AutoProcessor, AutoModelForCausalLM
45repo ="audarai/Audar-ASR-V1-Turbo"6proc = AutoProcessor.from_pretrained(repo, trust_remote_code=True)7model = AutoModelForCausalLM.from_pretrained(8 repo, trust_remote_code=True,9 dtype=torch.bfloat16, device_map="cuda:0",10).eval()1112SYSTEM ="فرّغ الكلام العربي التالي."# "Transcribe the following Arabic speech."13audio, _ = librosa.load("clip.wav", sr=16000, mono=True)1415conv =[{"role":"system","content": SYSTEM},16{"role":"user","content":[{"type":"audio"}]}]# audio placeholder (a list, not "<audio>")17text = proc.apply_chat_template(conv, tokenize=False, add_generation_prompt=True)18inputs = proc(text=text, audio=audio, sampling_rate=16000, return_tensors="pt").to(model.device)19inputs["input_features"]= inputs["input_features"].to(model.dtype)# features are fp32 -> cast to bf162021out = model.generate(**inputs, max_new_tokens=440, do_sample=False)22print(proc.batch_decode(out[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)[0].strip())
The self-contained modeling code targets transformers==4.57.6 (the version this checkpoint was
built and validated with). For version-independent, high-throughput serving, prefer vLLM — it
implements Qwen3-ASR natively (no custom code); see below.
File (repo root)
Approx. size
Notes
model.safetensors
~4.7 GB
Full bf16 weights (2,349,217,408 params)
config.json · *_audar_asr.py · __init__.py
—
Config + self-contained Qwen3-ASR modeling code
tokenizer files · preprocessor_config.json
—
Qwen3 tokenizer + Whisper-mel feature extractor
⚡ vLLM inference (GPU serving)
Turbo also runs on vLLM for high-throughput GPU serving with an
OpenAI-compatible API. vLLM implements the Qwen3-ASR architecture natively
(Qwen3ASRForConditionalGeneration + Qwen3ASRRealtimeGeneration) — no custom serving code: point vLLM at a
checkpoint and it exposes /v1/chat/completions, /v1/audio/transcriptions, and a realtime /v1/realtime
WebSocket.
vLLM serves quantized compressed-tensors checkpoints (not the GGUF files — those are for llama.cpp).
A vLLM-ready 4-bit (W4A16) build is provided in the vllm-w4a16/ folder:
Build
Folder
Size
Decoder
Audio encoder / lm_head / embeddings
Accuracy
W4A16
vllm-w4a16
~2.6 GB
INT4 (group-128)
BF16 (kept)
~+1 pp CER vs BF16
Only the language-model decoder is quantized; the audio encoder + projector stay BF16 (the projector's
ClippableLinear is numerically sensitive — the same rule as the GGUF mmproj), as do lm_head and the token
embeddings. An FP8 build (lossless vs BF16, ~3.3 GB) can be produced with the same recipe — see the note at
the end.
For full-precision GPU serving, point vLLM at the repo (the full bf16 root weights) instead of the 4-bit build — same native Qwen3-ASR support, no quantization.
1. Install (audio support required)
vLLM needs the audio extras (PyAV + librosa + soundfile) to decode audio; the stock image does not ship them:
dockerfile
1FROM vllm/vllm-openai:v0.24.02RUN pip install --no-cache-dir av librosa soundfile
docker build -t vllm-audio:0.24.0 .
(Or in a plain environment: pip install "vllm>=0.24" av librosa soundfile.)
vLLM auto-detects the compressed-tensors quantization (Marlin INT4 kernel). Weights + KV cache fit on any
≥12 GB GPU.
3. Transcribe
Turbo is prompt-steerable: the system message sets the task/language. For Arabic use
فرّغ الكلام العربي التالي.; steer other languages with the equivalent instruction. Send 16 kHz mono audio as
base64 input_audio and decode greedily (temperature: 0).
The OpenAI-style POST /v1/audio/transcriptions (multipart file upload) endpoint is also available for
Whisper-style clients.
4. Accuracy (FLEURS Arabic, greedy)
Character Error Rate vs the BF16 source — CER is the stable cross-precision metric for Arabic, where minor
و-segmentation differences inflate WER without changing the characters:
Build
AR CER
Δ vs BF16
BF16 source
2.46 %
—
W4A16 (this build)
3.73 %
+1.27
FP8 (optional)
2.46 %
+0.00 (lossless)
Leaderboard-grade full-test-set numbers are in the Benchmarks section above; 4-bit quantization
keeps them within ~1 pp CER (FP8 keeps them exactly).
Notes
Realtime streaming: vLLM also registers Qwen3ASRRealtimeGeneration, exposing an
OpenAI-Realtime-compatible /v1/realtime WebSocket; pair it with VAD/endpointing for stable incremental output.
Long audio: the audio encoder is a 30 s window; chunk longer inputs client-side.
Producing other precisions (needs the BF16 source weights): quantize the decoder Linears only via
llm-compressormodel_free_ptq, ignoring the audio tower,
lm_head, and embeddings — scheme="W4A16" (4-bit) or "FP8_DYNAMIC" (lossless),
ignore=["re:.*lm_head.*","re:.*embed_tokens.*","re:.*audio_tower.*"].
🎙️ Real-time streaming
Audar-ASR streams via LocalAgreement-2: as audio arrives the trailing window is re-decoded each hop
and a word is committed only once two consecutive decodes agree on it — giving stable, low-latency
incremental output over the GGUF runtime. Audar's production realtime engine serves the same policy over
an OpenAI-Realtime-compatible WebSocket with model-based endpointing and ≥64 concurrent streams on a
single A100-80GB.
🌍 Languages, dialects & tasks
Primary: Arabic — MSA and dialectal (Gulf/Emirati, Egyptian, Levantine, Maghrebi), plus
code-switched Arabic–English; emits dialect-faithful orthography from audio alone.
Also: English + 28 additional languages.
Task: transcription (audio → UTF-8 text), prompt-steerable for language and formatting.
Intended use & limitations
Intended use. Broadcast/media transcription, meeting & contact-center intelligence, voice agents,
captioning, and accessibility — cloud or on-prem.
Limitations.
Maghrebi / Moroccan Darija (Casablanca) remains the hardest condition (~63 % WER) for all systems.
Heavily code-switched telephony and low-SNR audio degrade accuracy relative to clean MSA.
Long-form audio can drift on very long recordings.
Not evaluated for, and must not be used for, covert speaker identification.
📜 License
Released under the AudarAI Community License v1.0 — research and limited commercial use for
qualifying Community Entities; enterprise / large-scale / MaaS use requires an AudarAI Enterprise
License. See
audarai.com/license/audarai-community-license-v1.0.
Citation
bibtex
1@misc{audar-asr-turbo-2026,
2 title = {Audar-ASR-V1: A Multilingual, Arabic-First Generative Speech Recognition Foundation Model},
3 author = {AudarAI},
4 year = {2026},
5 note = {Audar-ASR-V1-Turbo},
6 url = {https://github.com/AudarAI/Audar-ASR-V1/blob/main/report/Audar-ASR-V1-Technical-Report.pdf}
7}
About AudarAI
Leading Arabic-First Multilingual Audio Intelligence
AudarAI starts with Arabic — and expands to the world.
We are building advanced multilingual audio intelligence that helps individuals, enterprises, and
governments communicate across languages, cultures, and borders. By combining Arabic-first speech
technology with global multilingual AI, AudarAI transforms voice into understanding, interaction,
and connection.
Our work spans speech recognition, speech understanding, voice-enabled digital assistants,
human-computer interaction, and intelligent audio systems designed for real-world impact. From
empowering people to access technology in their native language to helping organizations
communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and
connected.
Arabic-first. Multilingual by design. Human-centered at heart.