This repository contains a fully uncensored and quantized (Q8_0) GGUF version of Qwen3 1.7B, designed for offline, local inference using llama.cpp and compatible runtimes.
By default, the model operates in thinking mode.
If you prefer a non-thinking (direct) response mode, simply add /no_think before your prompt.