Views
No views yet
o_proj 與所有 MoE 專家的 down_proj,共 2,973 個矩陣)正交化移除,
強度 0.8。Mamba(SSM) 的 out_proj 與 MTP 草稿頭保持原狀,以維持混合架構的連貫性。
| 量化 | 檔案大小 | 說明 |
|---|---|---|
| Q8_0 | 33 GB | 近乎無損 |
| Q6_K | 33 GB | 高品質(見下方註) |
| Q5_K_M | 26 GB | 品質與體積平衡 |
| Q4_K_M | 24 GB | 一般部署建議 |
| IQ4_XS | 18 GB | 以 importance matrix 量化 |
| IQ2_M | 18 GB | 最小可用,以 importance matrix 量化(見下方註) |
imatrix.dat(量化用的 importance matrix)。註(MoE 量化特性):本模型 hidden 維度為 2688,非 256 的整數倍,部分專家張量在量化時會回退到較高位元, 因此 Q8_0 與 Q6_K、IQ4_XS 與 IQ2_M 的體積相近。各版本皆可正常使用;追求最小體積請選 IQ4_XS 或 IQ2_M, 追求品質請選 Q8_0 或 Q6_K。
1llama-server -m Nemotron-3.5-Lightning-30B-A3B-Uncensored-xCloud-Q4_K_M.gguf \
2 --jinja -ngl 99 -c 8192max_tokens(建議 ≥ 2000),思考才收斂、答案才會出現;避免使用 medium_effort(思考更長且不收斂)。o_proj and every MoE expert's
down_proj; 2,973 matrices) at strength 0.8. The Mamba (SSM) out_proj and the MTP draft head are left
intact to preserve the coherence of the hybrid architecture.
| Quant | Size | Notes |
|---|---|---|
| Q8_0 | 33 GB | near-lossless |
| Q6_K | 33 GB | high quality (see note) |
| Q5_K_M | 26 GB | quality/size balance |
| Q4_K_M | 24 GB | recommended for deployment |
| IQ4_XS | 18 GB | importance-matrix quantized |
| IQ2_M | 18 GB | smallest usable, importance-matrix quantized (see note) |
imatrix.dat.Note (MoE quantization): this model's hidden size is 2688, not a multiple of 256, so some expert tensors fall back to higher bit-widths during quantization. As a result Q8_0/Q6_K and IQ4_XS/IQ2_M are close in size. All variants work normally; pick IQ4_XS/IQ2_M for the smallest footprint, Q8_0/Q6_K for the highest quality.
1llama-server -m Nemotron-3.5-Lightning-30B-A3B-Uncensored-xCloud-Q4_K_M.gguf \
2 --jinja -ngl 99 -c 8192max_tokens (>= 2000) so the thinking converges and the answer appears; avoid
medium_effort (longer, non-converging thinking).