Built for the Mistral AI Online Hackathon 2026 (W&B Fine-Tuning Track).
Two model variants available:
Evoxtral SFT — Best overall transcription accuracy (lowest WER)
Evoxtral RL — Best expressive tag accuracy (highest Tag F1)
What It Does
Standard ASR:
So I was thinking maybe we could try that new restaurant downtown. I mean if you're free this weekend.
Evoxtral:
[nervous] So... [stammers] I was thinking maybe we could... [clears throat] try that new restaurant downtown? [laughs nervously] I mean, if you're free this weekend?
SFT: LoRA finetuning on 808 synthetic audio samples with expressive tags (lr=2e-4, 3 epochs)
RL (RAFT): Rejection sampling — generate 4 completions per sample, score with rule-based reward (WER accuracy + Tag F1 - hallucination penalty), keep best, then SFT on curated data (lr=5e-5, 1 epoch)
Evaluated on 50 held-out test samples. Full benchmark (Evoxtral-Bench) with 7 metrics:
Core Metrics — Base vs SFT vs RL
Metric
Base Voxtral
Evoxtral SFT
Evoxtral RL
Best
WER
6.64%
4.47%
5.12%
SFT
CER
2.72%
1.23%
1.48%
SFT
Tag F1
22.0%
67.2%
69.4%
RL
Tag Precision
22.0%
67.4%
68.5%
RL
Tag Recall
22.0%
69.4%
72.7%
RL
Emphasis F1
42.0%
84.0%
86.0%
RL
Tag Hallucination
0.0%
19.3%
20.2%
SFT
SFT excels at raw transcription accuracy (best WER/CER). RL further improves expressive tag generation (+2.2% Tag F1, +3.3% Tag Recall, +2% Emphasis F1) at a small cost to WER.
Per-Tag F1 Breakdown (SFT → RL)
Tag
SFT F1
RL F1
Change
Support
[sighs]
1.000
1.000
—
9
[clears throat]
0.889
1.000
+12.5%
8
[gasps]
0.957
0.957
—
12
[pause]
0.885
0.902
+1.9%
25
[nervous]
0.800
0.846
+5.8%
13
[stammers]
0.889
0.842
-5.3%
8
[laughs]
0.800
0.815
+1.9%
12
[sad]
0.667
0.750
+12.4%
4
[whispers]
0.636
0.667
+4.9%
13
[crying]
0.750
0.571
-23.9%
5
[excited]
0.615
0.571
-7.2%
5
[shouts]
0.400
0.500
+25.0%
3
[calm]
0.200
0.400
+100%
6
[frustrated]
0.444
0.444
—
3
[angry]
0.667
0.667
—
2
[confused]
0.000
0.000
—
1
[scared]
0.000
0.000
—
1
RL improved 9 tags, kept 4 stable, and regressed 3. Biggest gains on [clears throat] (+12.5%), [calm] (+100%), [sad] (+12.4%), and [shouts] (+25%).