Views
No views yet
train split of the ru_llm_calibration dataset, and are intended to be run with llama.cpp.| Quantization type | File size | Quality | Recommendation |
|---|---|---|---|
| Q8_0 | ~8.05 GB | Virtually identical to FP16 | Best quality. Ideal for CPU inference when memory is not a constraint. |
| Q5_K_M | ~5.41 GB | Minimal degradation | Recommended balance. Excellent speed and quality, fits most consumer GPUs. |
| Q4_K_M | ~4.65 GB | Moderate degradation | "Golden standard". Best trade-off between size and quality. |
| IQ3_M | ~3.54 GB | Noticeable degradation | Maximum memory savings. Quality drops visibly; suited for highly constrained devices. |
test split of the Ru LLM Calibration dataset using the llama-perplexity utility. The original FP16 model served as the reference.| Metric | Q8_0 | Q5_K_M | Q4_K_M | IQ3_M |
|---|---|---|---|---|
| Mean PPL (Q) ↓ | 9.047 | 9.075 | 9.135 | 9.689 |
| PPL correlation ↑ | 99.97% | 99.87% | 99.69% | 98.64% |
| Mean KLD ↓ | 0.0020 | 0.0077 | 0.0174 | 0.0804 |
| Same top p ↑ | 96.71% | 94.36% | 92.16% | 84.58% |
↑ – higher is better; ↓ – lower is better
llama.cpp1# CLI
2./llama-cli -hf bond005/meno-lite-0.1-gguf -m meno-lite-0.1-Q4_K_M.gguf -p "Привет, как дела?"
3
4# Server with WebUI (default http://127.0.0.1:8080)
5./llama-server -hf bond005/meno-lite-0.1-gguf -m meno-lite-0.1-Q4_K_M.gguf --host 0.0.0.0 --port 8080llama.cpp documentation.