Views
No views yet
Q8_0, or F16 if you want the exact reference weights.TQ1_0, TQ2_0 and Q1_0 were built and tested but produced unusable output at this scale, so they are deliberately not published here.Q8_0 to IQ1_S saves you a few
tens of megabytes while degrading output substantially. The aggressive quants
are provided for completeness, not because they are a good trade.| File | Quant | Size | Notes |
|---|---|---|---|
Boris-250M-Instruct-F16.gguf | F16 | 481M | Reference. Unquantized conversion of the safetensors weights. |
Boris-250M-Instruct-BF16.gguf | BF16 | 481M | Reference, bfloat16. |
Boris-250M-Instruct-Q8_0.gguf | Q8_0 | 259M | Effectively lossless. Recommended. |
Boris-250M-Instruct-Q6_K.gguf | Q6_K | 244M | Near-lossless. |
Boris-250M-Instruct-Q5_K_M.gguf | Q5_K_M | 209M | Very good quality. |
Boris-250M-Instruct-Q5_K_S.gguf | Q5_K_S | 198M | |
Boris-250M-Instruct-Q5_1.gguf | Q5_1 | 202M | |
Boris-250M-Instruct-Q5_0.gguf | Q5_0 | 191M | |
Boris-250M-Instruct-Q4_K_M.gguf | Q4_K_M | 199M | Standard 4-bit default for larger models. |
Boris-250M-Instruct-Q4_K_S.gguf | Q4_K_S | 184M | |
Boris-250M-Instruct-Q4_1.gguf | Q4_1 | 179M | |
Boris-250M-Instruct-Q4_0.gguf | Q4_0 | 168M | |
Boris-250M-Instruct-IQ4_NL.gguf | IQ4_NL | 168M | |
Boris-250M-Instruct-IQ4_XS.gguf | IQ4_XS | 166M | |
Boris-250M-Instruct-Q3_K_L.gguf | Q3_K_L | 187M | |
Boris-250M-Instruct-Q3_K_M.gguf | Q3_K_M | 176M | |
Boris-250M-Instruct-Q3_K_S.gguf | Q3_K_S | 160M | |
Boris-250M-Instruct-IQ3_M.gguf | IQ3_M | 168M | |
Boris-250M-Instruct-IQ3_S.gguf | IQ3_S | 160M | |
Boris-250M-Instruct-IQ3_XS.gguf | IQ3_XS | 160M | |
Boris-250M-Instruct-IQ3_XXS.gguf | IQ3_XXS | 157M | |
Boris-250M-Instruct-Q2_K.gguf | Q2_K | 160M | |
Boris-250M-Instruct-Q2_K_S.gguf | Q2_K_S | 155M | |
Boris-250M-Instruct-IQ2_M.gguf | IQ2_M | 153M | |
Boris-250M-Instruct-IQ2_S.gguf | IQ2_S | 152M | |
Boris-250M-Instruct-IQ2_XS.gguf | IQ2_XS | 151M | |
Boris-250M-Instruct-IQ2_XXS.gguf | IQ2_XXS | 149M | |
Boris-250M-Instruct-IQ1_M.gguf | IQ1_M | 147M | |
Boris-250M-Instruct-IQ1_S.gguf | IQ1_S | 146M |
1# straight from the Hub
2llama-server -hf KSP-NMAI/Boris-250M-Instruct-GGUF:Q8_0 --jinja
3
4# or a local file
5llama-server -m Boris-250M-Instruct-Q8_0.gguf --jinja--jinja applies the
correct prompt format automatically.### Instruction:
{your instruction}
### Response: