Views
No views yet
imatrix) driven greedy allocation pipeline with class-specific quantization ladders, this build achieves better perplexity than the official Q6_K quantization while saving over 1.47 GB of disk and VRAM footprint.llama-perplexity on standard wikitext-2-raw (512 context chunks):| File | Size | BPW | PPL (Wikitext-2) | 8GB VRAM Usability |
|---|---|---|---|---|
| Ornith-1.5-9B-SHQ-6.5GB | 6.55 GB | ~6.1 | 8.9089 | Full GPU Offload (up to 16k context) |
| Ornith-1.5-9B-SHQ-5.4GB | 5.38 GB | ~5.0 | 9.2262 | Full GPU Offload (Long context 32k+) |
| Official Q6_K | 6.85 GB | 6.56 | 9.4833 | Tight (Risk of OOM on 8GB VRAM) |
| Official Q5_K_M | 6.40 GB | 5.50 | ~9.70+ | Fits in 8GB |
IQ4_NL, capturing semantic fidelity far better than traditional linear block quantizers.Q6_K, Q8_0, and F32), channeling aggressive compression only into less sensitive feed-forward layers.This project was developed to push the limits of modern architectures on consumer 8GB hardware. If you found this optimization useful and would like to support future quantizations and hardware upgrades, you can support me here: