Views
No views yet
cosyvoice.cpp) is MIT.Q6_K_S → Q6_KQ5_K_S → Q5_K_XXSQ4_K_S → Q4_K_XXSQ3_K_S → Q3_K_XXSQ2_K_S → removed (mostly noisy and unusable)| File | Quantization | Size (approx) | Notes |
|---|---|---|---|
CosyVoice3-2512_F32.gguf | F32 | 3.21 GiB | Highest precision, largest size |
CosyVoice3-2512_F16.gguf | F16 | 1.61 GiB | High-quality baseline |
CosyVoice3-2512_Q8_0.gguf | Q8_0 | 0.88 GiB | Near-F16 quality, smaller |
CosyVoice3-2512_Q6_K.gguf | Q6_K | 0.78 GiB | Formerly Q6_K_S; good quality/size balance |
CosyVoice3-2512_Q5_1.gguf | Q5_1 | 0.64 GiB | Better quality than Q5_0 |
CosyVoice3-2512_Q5_K_M.gguf | Q5_K_M | 0.64 GiB | New; maintains decent quality, good results |
CosyVoice3-2512_Q5_K_S.gguf | Q5_K_S | 0.61 GiB | New; maintains decent quality, good results |
CosyVoice3-2512_Q5_K_XXS.gguf | Q5_K_XXS | 0.61 GiB | Formerly Q5_K_S; further compressed |
CosyVoice3-2512_Q5_0.gguf | Q5_0 | 0.59 GiB | Compact mid-quality |
CosyVoice3-2512_Q4_K_M.gguf | Q4_K_M | 0.59 GiB | New; quality drops from Q5, occasional muffled audio, mostly acceptable |
CosyVoice3-2512_Q4_K_S.gguf | Q4_K_S | 0.56 GiB | New; lower quality than Q4_K_M |
CosyVoice3-2512_Q4_1.gguf | Q4_1 | 0.54 GiB | Smaller size, audible pronunciation artifacts |
CosyVoice3-2512_Q4_K_XXS.gguf | Q4_K_XXS | 0.54 GiB | Formerly Q4_K_S |
CosyVoice3-2512_Q4_0.gguf | Q4_0 | 0.49 GiB | Further quality degradation, often unclear speech |
CosyVoice3-2512_Q3_K_M.gguf | Q3_K_M | 0.49 GiB | New; muffled audio, slightly below but close to Q4_K_XXS |
CosyVoice3-2512_Q3_K_S.gguf | Q3_K_S | 0.46 GiB | New; lower quality than Q3_K_M |
CosyVoice3-2512_Q3_K_XXS.gguf | Q3_K_XXS | 0.44 GiB | Formerly Q3_K_S |
CosyVoice3-2512_Q2_K_L.gguf | Q2_K_L | 0.43 GiB | New; significant quality drop, quiet and muffled with noise |
CosyVoice3-2512_Q2_K.gguf | Q2_K | 0.42 GiB | New; severe artifacts and noise, but speech content can still be discerned |
Q8_0 (near-lossless listening quality in current tests, with much smaller size than F16)F16 and Q6_KQ5_1, Q5_K_M, Q5_K_S, Q5_0, Q5_K_XXSQ4_K_M, Q4_K_SQ4_1, Q4_K_XXS, Q4_0Q3_K_M, Q3_K_S, Q3_K_XXSQ2_K_L, Q2_KQ8_0: quality remains strong with little audible loss in typical samples.Q6_K: still sounds good and is often a practical choice.Q5_K_M / Q5_K_S (new): maintain decent quality, good results.Q5_K_XXS (formerly Q5_K_S): slightly higher than but close to Q4_K_M.Q5 family (Q5_1 / Q5_K_M / Q5_K_S / Q5_0 / Q5_K_XXS): generally usable with moderate degradation.Q4_K_M / Q4_K_S (new): quality drops compared to Q5, occasional muffled audio, but mostly acceptable.Q4 family (including Q4_1 / Q4_K_XXS / Q4_0): audible pronunciation artifacts, often muffled.Q3_K_M / Q3_K_S (new): further quality degradation, audio is muffled, slightly below but close to Q4_K_XXS.Q3_K_XXS (formerly Q3_K_S): aggressive quantization, limited quality.Q2_K_L (new): significant quality drop, quieter and muffled audio with noise.Q2_K (new): severe artifacts and noise, but speech content can still be discerned.Q2_K_Scosyvoice.cpp GGUF inference pipeline.1cosyvoice-cli \
2 --model CosyVoice3-2512_Q8_0.gguf \
3 --prompt-speech prompt_speech.gguf \
4 --text "Hello from CosyVoice" \
5 --output out.wav--prompt-speech expects a prompt-speech file in GGUF format (for example prompt_speech.gguf).cosyvoice.cpp repository:llama.cpp and modified for this project.cosyvoice.cpp): MIT