The compact build of NVIDIA's Parakeet TDT 0.6B v3, converted to GGUF for
transcribe.cpp. Same 25-language offline speech-to-text model and same architecture as the Q8_0 release, quantised harder.
Roughly two thirds the download of the Q8_0 build for a very small accuracy cost: 1.98% word error rate on LibriSpeech test-clean, against 1.94% for Q8_0. Does not stream and does not translate.
Quantization performed at the Faculty of Engineering, McMaster University.
The GGUF conversion this build is derived from was produced by
handy-computer, and the
weights here are a byte-for-byte copy of that file — the SHA-256 above matches
the upstream artifact.
Expects 16 kHz mono audio.
Licensed CC-BY-4.0. Attribution to NVIDIA is required when redistributing this model or its derivatives.