A q4_K_M GGUF of Qwen2.5-14B-Instruct with the dolly LoRA merged in, built with
llama.cpp. Part of a full open-model loop demo: fine-tune, merge, quantize, serve.
On a quant-ladder benchmark of this model (H100 SXM 80GB), q4_K_M was the fastest
at generation, about 126 tok/s on GPU and the fastest on CPU as well, while
costing only about 3.5 percent perplexity over q8_0 on wikitext-2. Smallest,
fastest, minimal quality loss. Full numbers and method:
https://github.com/ArgusForge/open-model-ops