This is an 8-bit quantized
MLX version of
Voxtral Mini 4B Realtime by Mistral AI, converted using
voxmlx.
This version was created for use with
Supervoxtral, enabling blazingly-fast realtime transcription on MacOS.
Voxtral Mini is a speech-to-text model that supports 13+ languages with sub-500ms latency. This version has been quantized to 8-bit precision for efficient inference on Apple Silicon using the MLX framework.