Call-grade speech recognition for 11 Indian languages.
By Tausif Iqbal. Built on NVIDIA NeMo FastConformer, fine-tuned from VAANI for PSTN, VoIP, and contact-center audio.
VaaniCall transcribes a phone call without a language code. It infers language and script, and it is trained through a 21-stage telephony pipeline so AMR, packet loss, and handset noise look like training data — not a domain shift.
All numbers are Word Error Rate (lower is better). Evaluations use 500 clips per language.
Overall WER: VaaniCall vs Whisper large-v3 on IndicVoices and FLEURS
Benchmark
Condition
VaaniCall
Whisper large-v3
Relative WER cut
IndicVoices
Clean
24.20%
84.07%
71%
IndicVoices
Telephonic
24.06%
84.50%
72%
Google FLEURS
Clean
23.52%
65.23%
64%
Whisper large-v3 still leads on English. VaaniCall is the Indic-language model.
VaaniCall WER stays flat from clean to telephonic audio
Benchmarks
vs Whisper large-v3 — IndicVoices, telephonic
Per-language telephonic WER vs Whisper large-v3
IndicVoices · telephonic · full table
Language
N
VaaniCall WER ↓
VaaniCall CER ↓
Whisper WER
Whisper CER
Bengali
500
15.41%
4.93%
73.87%
38.04%
English
500
13.26%
4.56%
7.80%
3.04%
Gujarati
500
21.46%
7.56%
61.48%
29.72%
Hindi
500
16.04%
6.87%
30.31%
15.61%
Kannada
500
38.01%
10.96%
98.03%
53.05%
Malayalam
500
40.70%
11.53%
138.83%
108.41%
Marathi
500
17.10%
6.30%
93.50%
47.21%
Odia
500
31.52%
10.40%
110.63%
97.36%
Punjabi
500
14.46%
6.03%
86.10%
54.26%
Tamil
500
33.82%
9.84%
80.72%
40.71%
Telugu
500
28.45%
8.63%
146.64%
97.12%
Overall
5500
24.06%
8.26%
84.50%
56.30%
IndicVoices · clean · full table
Language
N
VaaniCall WER ↓
VaaniCall CER ↓
Whisper WER
Whisper CER
Bengali
500
15.16%
4.96%
73.36%
36.81%
English
500
14.14%
4.67%
8.37%
3.12%
Gujarati
500
20.02%
6.68%
56.27%
25.96%
Hindi
500
16.45%
6.77%
30.71%
15.20%
Kannada
500
38.26%
11.02%
99.44%
52.16%
Malayalam
500
41.79%
12.11%
145.50%
114.81%
Marathi
500
16.76%
5.80%
92.38%
46.12%
Odia
500
31.85%
10.63%
107.96%
95.15%
Punjabi
500
15.71%
6.59%
89.11%
58.95%
Tamil
500
34.50%
10.22%
77.39%
38.11%
Telugu
500
28.39%
8.73%
141.52%
96.52%
Overall
5500
24.20%
8.31%
84.07%
56.30%
vs Whisper large-v3 — Google FLEURS (out of domain)
Held-out read speech. Odia is not in FLEURS.
FLEURS out-of-domain WER vs Whisper large-v3
FLEURS · clean · full table
Language
N
VaaniCall WER ↓
VaaniCall CER ↓
Whisper WER
Whisper CER
Bengali
500
26.00%
9.22%
77.60%
32.63%
English
500
15.16%
7.88%
4.47%
1.78%
Gujarati
500
24.94%
8.22%
54.99%
22.19%
Hindi
500
15.70%
6.18%
27.61%
9.63%
Kannada
500
25.80%
7.95%
72.85%
22.87%
Malayalam
500
37.55%
10.36%
130.76%
102.00%
Marathi
500
22.32%
7.18%
76.52%
22.61%
Punjabi
500
16.53%
6.06%
74.63%
35.88%
Tamil
500
32.78%
9.84%
47.73%
12.11%
Telugu
500
29.19%
9.99%
126.29%
74.64%
Overall
5000
23.52%
8.34%
65.23%
34.04%
Tamil CER is lower for Whisper (12.11%) despite a higher WER — typically an artifact of shorter predicted strings. VaaniCall still wins WER.
Per-language profile
VaaniCall WER by language on telephonic IndicVoices
Strongest Indic languages on phone audio: Punjabi 14.5% · Bengali 15.4% · Hindi 16.0% · Marathi 17.1%.
vs VAANI base — the telephony fine-tune
VaaniCall is a telephony specialist on top of VAANI. On phone-channel audio it improves 9 of 10 Indic languages. On clean audio the two models are statistically tied (24.37% vs 24.38% WER).
WER reduction versus VAANI base under telephonic conditions
vs VAANI base · telephonic
Language
N
VaaniCall WER ↓
VAANI WER
ΔWER
VaaniCall CER ↓
VAANI CER
ΔCER
Bengali
500
19.85%
21.70%
+1.85
7.74%
9.36%
+1.62
Gujarati
500
26.27%
28.97%
+2.70
10.72%
11.50%
+0.78
Hindi
500
19.65%
20.28%
+0.63
8.92%
9.22%
+0.30
Kannada
500
44.44%
47.03%
+2.59
14.93%
15.73%
+0.80
Malayalam
500
47.67%
48.51%
+0.84
16.46%
16.79%
+0.33
Marathi
500
23.65%
25.12%
+1.46
9.29%
10.02%
+0.73
Odia
500
35.93%
34.56%
−1.37
13.65%
13.28%
−0.37
Punjabi
500
20.26%
20.78%
+0.51
9.50%
9.71%
+0.21
Tamil
500
40.88%
41.84%
+0.96
14.69%
15.10%
+0.41
Telugu
500
32.17%
33.04%
+0.88
11.41%
11.75%
+0.34
Overall
5000
29.96%
31.04%
+1.08
11.96%
12.46%
+0.51
Δ = VAANI − VaaniCall. Positive Δ → VaaniCall is better. English excluded (not a VAANI fine-tune target).
vs VAANI base · clean
Language
N
VaaniCall WER
VAANI WER
ΔWER
VaaniCall CER
VAANI CER
ΔCER
Bengali
500
14.94%
15.07%
+0.13
5.01%
4.97%
−0.04
Gujarati
500
21.11%
21.93%
+0.82
7.38%
7.61%
+0.24
Hindi
500
16.47%
16.39%
−0.08
6.78%
6.75%
−0.03
Kannada
500
37.91%
38.78%
+0.87
10.99%
11.16%
+0.17
Malayalam
500
40.37%
40.63%
+0.26
11.46%
11.60%
+0.14
Marathi
500
17.05%
17.01%
−0.04
5.38%
5.33%
−0.05
Odia
500
31.18%
28.49%
−2.69
9.68%
9.15%
−0.53
Punjabi
500
14.17%
14.38%
+0.21
5.93%
5.86%
−0.07
Tamil
500
34.26%
34.49%
+0.24
9.82%
9.87%
+0.05
Telugu
500
27.12%
27.55%
+0.43
7.95%
8.07%
+0.12
Overall
5000
24.37%
24.38%
+0.01
8.17%
8.18%
+0.01
Results are statistically equivalent on clean audio — the fine-tune does not spend general accuracy to buy telephony robustness.
vs AI4Bharat IndicConformer
IndicConformer is a strong general-purpose Indic ASR and still leads this comparison. VaaniCall is optimized for call-channel robustness and zero-shot language ID, not for matching a dedicated per-language conformer on clean read speech.
Condition
VaaniCall
IndicConformer
Gap
Telephonic
29.38%
26.11%
−3.26 pp
Clean
24.56%
20.63%
−3.93 pp
Largest telephonic gaps: Malayalam −5.61 pp, Telugu −5.49 pp. English excluded.
vs IndicConformer · telephonic
Language
N
VaaniCall WER
IndicConformer WER ↓
ΔWER
VaaniCall CER
IndicConformer CER ↓
ΔCER
Bengali
500
18.99%
16.81%
−2.18
7.29%
6.65%
−0.63
Gujarati
500
25.31%
22.48%
−2.83
9.89%
9.46%
−0.43
Hindi
500
19.61%
17.27%
−2.34
9.25%
7.86%
−1.40
Kannada
500
44.50%
43.22%
−1.28
15.24%
13.92%
−1.33
Malayalam
500
45.91%
40.30%
−5.61
15.74%
12.75%
−2.99
Marathi
500
22.31%
18.78%
−3.53
8.27%
7.09%
−1.18
Odia
500
36.38%
32.34%
−4.04
13.86%
12.42%
−1.44
Punjabi
500
19.16%
16.48%
−2.68
8.75%
8.74%
−0.01
Tamil
500
38.82%
35.30%
−3.52
13.53%
11.34%
−2.18
Telugu
500
34.36%
28.87%
−5.49
12.52%
10.10%
−2.41
Overall
5000
29.38%
26.11%
−3.26
11.62%
10.15%
−1.47
vs IndicConformer · clean
Language
N
VaaniCall WER
IndicConformer WER ↓
ΔWER
VaaniCall CER
IndicConformer CER ↓
ΔCER
Bengali
500
15.42%
12.79%
−2.64
4.96%
4.00%
−0.95
Gujarati
500
20.37%
15.24%
−5.13
7.09%
5.11%
−1.98
Hindi
500
15.97%
14.23%
−1.74
6.58%
5.68%
−0.89
Kannada
500
37.67%
33.94%
−3.73
11.10%
9.40%
−1.70
Malayalam
500
41.67%
35.23%
−6.43
12.06%
9.67%
−2.39
Marathi
500
16.08%
13.02%
−3.05
5.26%
4.28%
−0.99
Odia
500
30.82%
25.86%
−4.96
9.86%
8.01%
−1.85
Punjabi
500
15.00%
10.65%
−4.35
6.48%
4.21%
−2.27
Tamil
500
34.76%
30.35%
−4.41
10.31%
8.52%
−1.79
Telugu
500
27.90%
24.43%
−3.47
8.49%
6.92%
−1.57
Overall
5000
24.56%
20.63%
−3.93
8.40%
6.74%
−1.66
How it works
Phone audio into FastConformer encoder, RNNT decoder, transcript with no language ID
Training
Two-phase fine-tune from the VAANI FastConformer checkpoint.
Phase 1 language alignment, phase 2 acoustic refinement
Language alignment — encoder frozen; decoder learns script and language mapping.
Acoustic refinement — last five encoder layers plus the full decoder unfrozen for telephonic acoustics.
Data
Trained on AI4Bharat IndicVoices and VAANI: 50,000 train / 5,000 val / 5,000 test clips per language (mean duration ~18 s).
Telephony augmentation
Every training batch can pass through a 21-stage, GPU-accelerated call-channel simulator.
Talker, handset, codec, network, and line stages of the augmentation pipeline
All 21 stages
#
Stage
Prob
What it does
0
Identity setup
100%
Tier: Premium / Standard / Legacy. Locks mic profile and codec behavior.
1
Formant shift
25%
Vocal-tract warp (0.95–1.05×) so the model does not memorize speakers.
2
Speed perturbation
30%
Tempo + pitch together (0.9–1.1×).
3
Room IR
50%
GPU conv1d reverb, partial RMS match, up to 20 ms pre-delay.
No language flag. 16 kHz mono WAV is the happy path; typical telephony (8 kHz, μ-law / AMR) is the training domain.
Intended use
Inbound / outbound contact-center transcription in Indic languages
VoIP and PSTN call analytics
Multilingual IVR and voice-bot post-processing
Offline batch transcription of call recordings
Not intended for: medical or legal dictation, real-time emergency dispatch, or English-only broadcast ASR (use Whisper or a dedicated English model).
Limitations
English trails Whisper large-v3 (expected: this checkpoint is Indic-first).
Malayalam, Kannada, Tamil remain the hardest languages (WER 34–41% on IndicVoices).
Odia slightly regresses vs the VAANI base under telephony (−1.37 pp).
IndicConformer is still stronger as a general Indic ASR.
Code-mixed utterances and heavy music-on-hold are not separately benchmarked.
Evaluations are 500 clips / language; treat per-language gaps of <1 pp as noise.
Citation
If you use VaaniCall, please cite this model and VAANI, the base the fine-tune starts from:
bibtex
1@misc{iqbal2026vaanicall,
2 title={VaaniCall: Multilingual Telephony ASR for Indic Languages},
3 author={Iqbal, Tausif},
4 year={2026},
5 howpublished={Hugging Face},
6 url={https://huggingface.co/TieIncred/VaaniCall}
7}
bibtex
1@misc{pulikodan2026vaanicapturinglanguagelandscape,
2 title={VAANI: Capturing the language landscape for an inclusive digital India},
3 author={Sujith Pulikodan and Abhayjeet Singh and Agneedh Basu and Nihar Desai and Pavan Kumar J and Pranav D Bhat and Raghu Dharmaraju and Ritika Gupta and Sathvik Udupa and Saurabh Kumar and Sumit Sharma and Vaibhav Vishwakarma and Visruth Sanka and Dinesh Tewari and Harsh Dhand and Amrita Kamat and Sukhwinder Singh and Shikhar Vashishth and Partha Talukdar and Raj Acharya and Prasanta Kumar Ghosh},
4 year={2026},
5 eprint={2603.28714},
6 archivePrefix={arXiv},
7 primaryClass={eess.AS},
8 url={https://arxiv.org/abs/2603.28714}
9}
Also cite IndicVoices when reporting numbers on that test set.