Views
No views yet
Q8_0, or Q6_K if you want the same quality a little smaller.d_model is 768 and the QKV/FFN rows are
2304/3072, all divisible by 256, so unlike the 75M sibling every tensor here
really is quantized at its nominal bit width — the k-quants and i-quants behave
as intended.Q6_K (~32 MB) in nearly every file here, which is most of the
50–60 MB you see at the bottom of the table.Q8_0 (130 MB) to IQ1_S (50 MB) saves 80 MB while more than tripling
perplexity, and the ternary files are broken outright. Below roughly Q3_K_M
the curve turns sharply and is not worth walking down unless you are curious.| File | Quant | Size | PPL | Notes |
|---|---|---|---|---|
Boris-1.3-125M-Instruct-F16.gguf | F16 | 241M | 55.10 | Reference. Unquantized conversion of the safetensors weights. |
Boris-1.3-125M-Instruct-BF16.gguf | BF16 | 241M | 55.12 | Reference, bfloat16. |
Boris-1.3-125M-Instruct-Q8_0.gguf | Q8_0 | 130M | 55.08 | Effectively lossless. Recommended. |
Boris-1.3-125M-Instruct-Q6_K.gguf | Q6_K | 102M | 55.16 | Near-lossless. |
Boris-1.3-125M-Instruct-Q5_K_M.gguf | Q5_K_M | 94M | 55.61 | Very good quality. |
Boris-1.3-125M-Instruct-Q5_K_S.gguf | Q5_K_S | 91M | 55.54 | |
Boris-1.3-125M-Instruct-Q5_1.gguf | Q5_1 | 96M | 55.51 | |
Boris-1.3-125M-Instruct-Q5_0.gguf | Q5_0 | 91M | 55.60 | |
Boris-1.3-125M-Instruct-Q4_K_M.gguf | Q4_K_M | 87M | 56.05 | Standard 4-bit default for larger models. |
Boris-1.3-125M-Instruct-Q4_K_S.gguf | Q4_K_S | 82M | 56.28 | |
Boris-1.3-125M-Instruct-Q4_1.gguf | Q4_1 | 86M | 56.80 | |
Boris-1.3-125M-Instruct-Q4_0.gguf | Q4_0 | 81M | 56.94 | |
Boris-1.3-125M-Instruct-IQ4_NL.gguf | IQ4_NL | 81M | 56.22 | Best of the 4-bit-and-under group per byte. |
Boris-1.3-125M-Instruct-IQ4_XS.gguf | IQ4_XS | 78M | 56.27 | |
Boris-1.3-125M-Instruct-Q3_K_L.gguf | Q3_K_L | 84M | 57.61 | |
Boris-1.3-125M-Instruct-Q3_K_M.gguf | Q3_K_M | 78M | 58.24 | Reasonable floor for everyday use. |
Boris-1.3-125M-Instruct-Q3_K_S.gguf | Q3_K_S | 70M | 61.97 | |
Boris-1.3-125M-Instruct-IQ3_M.gguf | IQ3_M | 74M | 58.07 | |
Boris-1.3-125M-Instruct-IQ3_S.gguf | IQ3_S | 70M | 60.44 | |
Boris-1.3-125M-Instruct-IQ3_XS.gguf | IQ3_XS | 69M | 61.27 | |
Boris-1.3-125M-Instruct-IQ3_XXS.gguf | IQ3_XXS | 62M | 65.56 | |
Boris-1.3-125M-Instruct-Q2_K.gguf | Q2_K | 67M | 68.54 | |
Boris-1.3-125M-Instruct-Q2_K_S.gguf | Q2_K_S | 62M | 83.13 | Starts looping on open-ended prompts. |
Boris-1.3-125M-Instruct-IQ2_M.gguf | IQ2_M | 60M | 68.66 | Best 2-bit option; matches Q2_K 7M smaller. |
Boris-1.3-125M-Instruct-IQ2_S.gguf | IQ2_S | 58M | 74.79 | |
Boris-1.3-125M-Instruct-IQ2_XS.gguf | IQ2_XS | 55M | 92.54 | Fluent but confabulates freely. |
Boris-1.3-125M-Instruct-IQ2_XXS.gguf | IQ2_XXS | 53M | 107.47 | Sometimes emits an empty answer. |
Boris-1.3-125M-Instruct-IQ1_M.gguf | IQ1_M | 51M | 153.38 | Still clean English, but largely unrelated to the question. |
Boris-1.3-125M-Instruct-IQ1_S.gguf | IQ1_S | 50M | 179.37 | |
Boris-1.3-125M-Instruct-TQ2_0.gguf | TQ2_0 | 56M | 8903.62 | Ternary; degenerate token soup. Curiosity only. |
Boris-1.3-125M-Instruct-TQ1_0.gguf | TQ1_0 | 52M | 8903.62 | Ternary; as above, and gives the same greedy output as TQ2_0. |
boris-1.3-125m-instruct.imatrix. Q8_0 and the legacy
Q4_0/Q4_1/Q5_0/Q5_1 are plain round-to-nearest. SHA256SUMS covers every file.b9013 (e48034dfc).1# straight from the Hub
2llama-server -hf KSP-NMAI/Boris-1.3-125M-Instruct-GGUF:Q8_0 --jinja
3
4# or a local file
5llama-server -m Boris-1.3-125M-Instruct-Q8_0.gguf --jinja--jinja applies the correct
prompt format automatically.<|endoftext|> (token 0), which is also the stop token.### System:
{optional system message}
### User:
{your message}
### Assistant:
{reply}<|endoftext|>
### User:
{next message}
### Assistant:arch=gpt2 with the chat template
present, and to complete chat-formatted prompts with output that decodes as strict
UTF-8 containing zero U+FFFD replacement characters — no mojibake at any quant
level, including the ternary files. The GGUF vocab was diffed against the source
tokenizer.json across accented Latin, Cyrillic, CJK, Hangul, Arabic, Hebrew,
emoji (including ZWJ sequences), zero-width and whitespace-run samples: identical
token ids and byte-exact detokenization.