He provides only q4 GGUF, but it's too large for my pc(I only have 12gb vram). So I quantized it so that there're also IQ2_M and IQ3_M GGUF, which both can run on machines with lower vram
plus: if you use runpod to quantize or train some models, don't train it directly in network volume. It's extremely instable and my quantizing script stopped for no reason every a few minutes
I've wasted at least 15$ for that.