Views
No views yet
Q8_0, or F16 if you want the exact reference weights.⚠️ Not recommended:IQ1_S. Included for completeness only.
TQ1_0, TQ2_0 and Q1_0 were built and tested but produced unusable output at this scale, so they are deliberately not published here.Q8_0 to IQ1_S saves you a few
tens of megabytes while degrading output substantially. The aggressive quants
are provided for completeness, not because they are a good trade.| File | Quant | Size | Notes |
|---|---|---|---|
Boris-125M-Instruct-F16.gguf | F16 | 240M | Reference. Unquantized conversion of the safetensors weights. |
Boris-125M-Instruct-BF16.gguf | BF16 | 240M | Reference, bfloat16. |
Boris-125M-Instruct-Q8_0.gguf | Q8_0 | 130M | Effectively lossless. Recommended. |
Boris-125M-Instruct-Q6_K.gguf | Q6_K | 101M | Near-lossless. |
Boris-125M-Instruct-Q5_K_M.gguf | Q5_K_M | 94M | Very good quality. |
Boris-125M-Instruct-Q5_K_S.gguf | Q5_K_S | 91M | |
Boris-125M-Instruct-Q5_1.gguf | Q5_1 | 96M | |
Boris-125M-Instruct-Q5_0.gguf | Q5_0 | 91M | |
Boris-125M-Instruct-Q4_K_M.gguf | Q4_K_M | 87M | Standard 4-bit default for larger models. |
Boris-125M-Instruct-Q4_K_S.gguf | Q4_K_S | 82M | |
Boris-125M-Instruct-Q4_1.gguf | Q4_1 | 86M | |
Boris-125M-Instruct-Q4_0.gguf | Q4_0 | 81M | |
Boris-125M-Instruct-IQ4_NL.gguf | IQ4_NL | 80M | |
Boris-125M-Instruct-IQ4_XS.gguf | IQ4_XS | 78M | |
Boris-125M-Instruct-Q3_K_L.gguf | Q3_K_L | 84M | |
Boris-125M-Instruct-Q3_K_M.gguf | Q3_K_M | 77M | |
Boris-125M-Instruct-Q3_K_S.gguf | Q3_K_S | 70M | |
Boris-125M-Instruct-IQ3_M.gguf | IQ3_M | 74M | |
Boris-125M-Instruct-IQ3_S.gguf | IQ3_S | 70M | |
Boris-125M-Instruct-IQ3_XS.gguf | IQ3_XS | 69M | |
Boris-125M-Instruct-IQ3_XXS.gguf | IQ3_XXS | 61M | |
Boris-125M-Instruct-Q2_K.gguf | Q2_K | 67M | |
Boris-125M-Instruct-Q2_K_S.gguf | Q2_K_S | 62M | |
Boris-125M-Instruct-IQ2_M.gguf | IQ2_M | 59M | |
Boris-125M-Instruct-IQ2_S.gguf | IQ2_S | 58M | |
Boris-125M-Instruct-IQ2_XS.gguf | IQ2_XS | 54M | |
Boris-125M-Instruct-IQ2_XXS.gguf | IQ2_XXS | 52M | |
Boris-125M-Instruct-IQ1_M.gguf | IQ1_M | 50M | |
Boris-125M-Instruct-IQ1_S.gguf | IQ1_S | 49M | ⚠️ Broken — incoherent on a greedy smoke test. |
1# straight from the Hub
2llama-server -hf KSP-NMAI/Boris-125M-Instruct-GGUF:Q8_0 --jinja
3
4# or a local file
5llama-server -m Boris-125M-Instruct-Q8_0.gguf --jinja--jinja applies the
correct prompt format automatically.### Instruction:
{your instruction}
### Response: