GGUF weights for omnivoice.cpp,
a C++17/GGML port of OmniVoice (k2-fsa/OmniVoice). Multilingual zero
shot TTS, 646 languages, 24 kHz mono. Runs on CPU, CUDA, ROCm, Metal,
Vulkan.
Set GGML_BACKEND to force a device, otherwise the runtime picks the
best one available.
value
target
CUDA0
NVIDIA GPU, fastest path on Ada / Blackwell
Vulkan0
Cross vendor GPU (AMD / Intel / NVIDIA)
Metal
Apple Silicon GPU
CPU
CPU fallback, x86 variant auto selected
Quantization policy
Tokenizer GGUFs are not uniform quants. Three categories get a
dedicated treatment :
tensor
dtype across all variants
RVQ codebooks, fc, fc2, project_in / project_out
F32
Snake activation alpha
F32
Conv kernels with non alignable rows (K=7,3,1)
F16 (in Q* variants)
Same fallback as llama.cpp tensor_type_fallback : F16 has no block
size and matches the runtime target dtype on every backend. The base
LM (Qwen3 0.6B, hidden = 1024) has all dimensions divisible by 256 so
the fallback never triggers, the LM follows standard llama.cpp K-quant
across variants.
License
Upstream model : OmniVoice by Xiaomi / k2-fsa, Apache 2.0
Audio codec : Higgs Audio v2 (bosonai/higgs-audio-v2-tokenizer), Apache 2.0
GGUF tooling : omnivoice.cpp, MIT