Views
No views yet
GLM-4.5-KLD-8192-ref-logits-ed-combined-all-micro-Q8_0.bin and combined_all_micro.txtglm-4.5-quantization-output.csv.[Default Type]-[FFN_UP]-[FFN_GATE]-[FFN_DOWN], eg: Q6_K-IQ2_S-IQ2_S-IQ3_S. This means:plots-glm-4.5-8192 folder.

llama.cpp and ik_llama.cpp folders for convenience (though, ik_llama.cpp is capable of running the quants in llama.cpp, but the opposite is not true).llama.cpp imatrix Quantizations of zai-org/GLM-4.5llama.cpp or kobold.cpp like normal.ik_llama.cpp imatrix Quantizations of zai-org/GLM-4.5ik_llama.cpp can also run your existing GGUFs from bartowski, unsloth, mradermacher, etc if you want to try it out before downloading my quants.1# 93 Repeating Layers [0-92]
2
3# Attention
4blk\..*\.attn_q.*=iq4_k
5blk\..*\.attn_k.*=iq6_k
6blk\..*\.attn_v.*=iq6_k
7blk\..*\.attn_output.*=iq5_ks
8
9# First 3 Dense Layers [0-2]
10blk\..*\.ffn_down\.weight=iq4_ks
11blk\..*\.ffn_(gate|up)\.weight=iq3_ks
12
13# Shared Expert Layers [3-92]
14blk\..*\.ffn_down_shexp\.weight=iq6_k
15blk\..*\.ffn_(gate|up)_shexp\.weight=iq6_k
16
17# Routed Experts Layers [3-92]
18blk\..*\.ffn_down_exps\.weight=iq3_kt
19blk\..*\.ffn_(gate|up)_exps\.weight=iq2_kt
20
21# NextN MTP Layer [92]
22blk\..*\.nextn\.embed_tokens\.weight=iq4_k
23blk\..*\.nextn\.shared_head_head\.weight=iq6_k
24blk\..*\.nextn\.eh_proj\.weight=iq6_k
25
26# Non-Repeating Layers
27token_embd\.weight=iq4_k
28output\.weight=iq6_k1# 93 Repeating Layers [0-92]
2
3# Attention
4blk\.(0|1|2)\.attn_q.*=q8_0
5blk\.(0|1|2)\.attn_k.*=q8_0
6blk\.(0|1|2)\.attn_v.*=q8_0
7blk\.(0|1|2)\.attn_output.*=q8_0
8
9blk\..*\.attn_q.*=iq6_k
10blk\..*\.attn_k.*=iq6_k
11blk\..*\.attn_v.*=iq6_k
12blk\..*\.attn_output.*=iq6_k
13
14# First 3 Dense Layers [0-2]
15blk\..*\.ffn_down\.weight=iq5_ks
16blk\..*\.ffn_(gate|up)\.weight=iq4_ks
17
18# Shared Expert Layers [3-92]
19blk\..*\.ffn_down_shexp\.weight=q8_0
20blk\..*\.ffn_(gate|up)_shexp\.weight=q8_0
21
22# Routed Experts Layers [3-92]
23blk\..*\.ffn_down_exps\.weight=iq4_ks
24blk\..*\.ffn_(gate|up)_exps\.weight=iq4_kss
25
26# NextN MTP Layer [92]
27blk\..*\.nextn\.embed_tokens\.weight=iq5_ks
28blk\..*\.nextn\.shared_head_head\.weight=iq5_ks
29blk\..*\.nextn\.eh_proj\.weight=q8_0
30
31# Non-Repeating Layers
32token_embd\.weight=iq4_k
33output\.weight=iq6_k1Default quant level @ Q8_0
2
3# Shared Expert Layers [3-92]
4blk\..*\.ffn_down_shexp\.weight=q8_0
5blk\..*\.ffn_(gate|up)_shexp\.weight=q8_0
6
7# Routed Experts Layers [3-92]
8blk\..*\.ffn_up_exps\.weight=iq4_ks
9blk\..*\.ffn_gate_exps\.weight=iq4_ks
10blk\..*\.ffn_down_exps\.weight=iq5_ks1# 93 Repeating Layers [0-92]
2
3# Attention
4blk\.(0|1|2)\.attn_q.*=q8_0
5blk\.(0|1|2)\.attn_k.*=q8_0
6blk\.(0|1|2)\.attn_v.*=q8_0
7blk\.(0|1|2)\.attn_output.*=q8_0
8
9blk\..*\.attn_q.*=iq5_k
10blk\..*\.attn_k.*=iq5_k
11blk\..*\.attn_v.*=iq5_k
12blk\..*\.attn_output.*=iq5_k
13
14# First 3 Dense Layers [0-2]
15blk\..*\.ffn_down\.weight=q8_0
16blk\..*\.ffn_(gate|up)\.weight=q8_0
17
18# Shared Expert Layers [3-92]
19blk\..*\.ffn_down_shexp\.weight=q8_0
20blk\..*\.ffn_(gate|up)_shexp\.weight=q8_0
21
22# Routed Experts Layers [3-92]
23blk\..*\.ffn_down_exps\.weight=iq5_k
24blk\..*\.ffn_(gate|up)_exps\.weight=iq4_k
25
26# NextN MTP Layer [92]
27blk\..*\.nextn\.embed_tokens\.weight=iq5_k
28blk\..*\.nextn\.shared_head_head\.weight=iq5_k
29blk\..*\.nextn\.eh_proj\.weight=q8_0
30
31# Non-Repeating Layers
32token_embd\.weight=q8_0
33output\.weight=q8_0