Views
No views yet
Q8_0, or F16 if you want the exact reference weights.Q8_0 to IQ1_S saves you a few
tens of megabytes while degrading output substantially. The aggressive quants
are provided for completeness, not because they are a good trade.| File | Quant | Size | Notes |
|---|---|---|---|
Boris-75M-Instruct-F16.gguf | F16 | 150M | Reference. Unquantized conversion of the safetensors weights. |
Boris-75M-Instruct-BF16.gguf | BF16 | 150M | Reference, bfloat16. |
Boris-75M-Instruct-Q8_0.gguf | Q8_0 | 82M | Effectively lossless. Recommended. |
Boris-75M-Instruct-Q6_K.gguf | Q6_K | 78M | Near-lossless. |
Boris-75M-Instruct-Q5_K_M.gguf | Q5_K_M | 69M | Very good quality. |
Boris-75M-Instruct-Q5_K_S.gguf | Q5_K_S | 66M | |
Boris-75M-Instruct-Q5_1.gguf | Q5_1 | 67M | |
Boris-75M-Instruct-Q5_0.gguf | Q5_0 | 65M | |
Boris-75M-Instruct-Q4_K_M.gguf | Q4_K_M | 67M | Standard 4-bit default for larger models. |
Boris-75M-Instruct-Q4_K_S.gguf | Q4_K_S | 63M | |
Boris-75M-Instruct-Q4_1.gguf | Q4_1 | 62M | |
Boris-75M-Instruct-Q4_0.gguf | Q4_0 | 59M | |
Boris-75M-Instruct-IQ4_NL.gguf | IQ4_NL | 59M | |
Boris-75M-Instruct-IQ4_XS.gguf | IQ4_XS | 58M | |
Boris-75M-Instruct-Q3_K_L.gguf | Q3_K_L | 64M | |
Boris-75M-Instruct-Q3_K_M.gguf | Q3_K_M | 61M | |
Boris-75M-Instruct-Q3_K_S.gguf | Q3_K_S | 57M | |
Boris-75M-Instruct-IQ3_M.gguf | IQ3_M | 59M | |
Boris-75M-Instruct-IQ3_S.gguf | IQ3_S | 57M | |
Boris-75M-Instruct-IQ3_XS.gguf | IQ3_XS | 57M | |
Boris-75M-Instruct-IQ3_XXS.gguf | IQ3_XXS | 56M | |
Boris-75M-Instruct-Q2_K.gguf | Q2_K | 57M | |
Boris-75M-Instruct-Q2_K_S.gguf | Q2_K_S | 56M | |
Boris-75M-Instruct-IQ2_M.gguf | IQ2_M | 55M | |
Boris-75M-Instruct-IQ2_S.gguf | IQ2_S | 55M | |
Boris-75M-Instruct-IQ2_XS.gguf | IQ2_XS | 55M | |
Boris-75M-Instruct-IQ2_XXS.gguf | IQ2_XXS | 54M | |
Boris-75M-Instruct-IQ1_M.gguf | IQ1_M | 54M | |
Boris-75M-Instruct-IQ1_S.gguf | IQ1_S | 53M | |
Boris-75M-Instruct-TQ2_0.gguf | TQ2_0 | 54M | |
Boris-75M-Instruct-TQ1_0.gguf | TQ1_0 | 53M |
1# straight from the Hub
2llama-server -hf KSP-NMAI/Boris-75M-Instruct-GGUF:Q8_0 --jinja
3
4# or a local file
5llama-server -m Boris-75M-Instruct-Q8_0.gguf --jinja--jinja applies the
correct prompt format automatically.### Instruction:
{your instruction}
### Response: