Qwen3 Coder 30B A3B Instruct gguf
Make sure you have enough ram/gpu to run. On the right of model card, you may see the size of each quantized models.
Work with openclaw
It is recommended to be used with openclaw to have a personal secretary. please set context length as 64K to 256K to be used locally. Local api provided by this model will not expose your secret to the internet.
Use the model in ollama
First download and install ollama.
Command
in windows command line, or in terminal in ubuntu, type:
ollama run hf.co/John1604/Qwen3-Coder-30B-A3B-Instruct-gguf:q5_k_m
(q5_k_m is the model quant type, q5_k_s, q4_k_m, ..., can also be used)
C:\Users\developer>ollama run hf.co/John1604/Qwen3-Coder-30B-A3B-Instruct-gguf:q5_k_m
Use the model in LM Studio
download and install LM Studio
Discover models
In the LM Studio, click "Discover" icon. "Mission Control" popup window will be displayed.
In the "Mission Control" search bar, type "John1604/Qwen3-Coder-30B-A3B-Instruct-gguf" and check "GGUF", the model should be found.
Download the model.
You may choose quantized model.
Load the model.
Ask questions.
quantized models
Type Bits Quality Description Q2_K 2-bit 🟥 Low Minimal footprint; only for tests Q3_K_S 3-bit 🟧 Low “Small” variant (less accurate) Q3_K_M 3-bit 🟧 Low–Med “Medium” variant Q4_K_S 4-bit 🟨 Med Small, faster, slightly less quality Q4_K_M 4-bit 🟩 Med–High “Medium” — best 4-bit balance Q5_K_S 5-bit 🟩 High Slightly smaller than Q5_K_M Q5_K_M 5-bit 🟩🟩 High Excellent general-purpose quant Q6_K 6-bit 🟩🟩🟩 Very High Almost FP16 quality, larger size Q8_0 8-bit 🟩🟩🟩🟩 Near-lossless baseline