Optimized with my branch's custom auto-tensor-type, custom-made recipes for 2.91, 3.74 and 4.34 bpw
(element-gamma=0.25, which beats the stock hand-tuned recipes at matched BPW).
Since HF doesn't recognize custom bpw tags, I've tagged them with:
Q2_K: 2.91bpw
IQ3_M: 3.74bpw
IQ4_XS: 4.34bpw
Note that the quant types are only aliases for the size and do not correspond to the actual quant types used.