Provided for conversion/experimentation; not independently benchmarked
Evaluation
The public weights are unchanged. These September 2026 results come from the
current documented runtime recipes on the same frozen 1,513-clip, 7.18-hour
medical benchmark and scorer. No dictionary, custom vocabulary, contextual
bias, reference-aware selection, or transcript correction was used.
Runtime artifact
Platform
WER
M-WER
Drug M-WER
Medical recall
NeMo canonical
NVIDIA CUDA (L4, BF16)
6.54%
2.23%
4.75%
97.77%
MLX q8
Apple Silicon (M4 Max)
6.65%
2.12%
4.52%
97.88%
GGUF q8_0
Linux/Windows CPU
7.10%
2.16%
4.30%
97.84%
The CPU path is the portable fallback. It recorded the lowest occurrence-based
medical and drug error counts in this draw, but paired tests do not establish
it as clinically better than GPU or MLX. Use CUDA for the best WER and
throughput, or MLX q8 on Apple Silicon.
The previous model-card figures were produced by an older runtime/scorer path.
The numbers above supersede them as runtime results; the GGUF weights did not
change.
Omi Med STT v1 is speech-to-text only. It is not a diagnostic, triage,
prescribing, or clinical decision model, and it is not clinically validated.
Transcripts must be reviewed before any clinical use.