This is a mixed quantization of
DarkArtsForge/Zepar-24B-v1 leveraging ik_llama.cpp SOTA quantization to compete with IQ4_XS mainline quants.
F16 GGUF and imatrix are avaliable
here.
iq4_ks: ffn_gate, ffn_up, ffn_down, token_embd
iq5_ks: attn_q, attn_v, attn_k, attn_output
iq6_k: output
imatrix created based on calibration data by
bartowski
I wanted an "IQ4_XS but better" with ik_llama for personal use. I made it. WYSIWYG.
MistralAI - the beloved base model(s).
ikawrakow and contributors of ik_llama.cpp - I probably misused your wonderful creation.
DarkArtsForge/Naphula - for the effort of designing and producing the merge. And for reviving my interest in Mistral 24b.
+++ Everyone whose finetunes were included in the merge!
bartowski - for the calibration data + the myriad of quants we all benefit from.
imatrix generated on an intermediary q8_0 quant. I don't eat RAM for breakfast.
My only contribution is compute. This is neither my merge, nor my calibration data. Have fun.