1# llama.cpp (server) — tool-calling needs the recovery shim, see below
2llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-GGUF:Q4_K_M --jinja --ctx-size 16384
3
4# Ollama
5ollama run hf.co/tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-GGUF:Q4_K_M
Sizes and a one-click loader are in the file browser / Quantizations widget above;
the note says which quant to reach for.
How this model was built — technique chain, training mix, and the exact knobs/pins,
so the result is reproducible without any of our tooling.
Something not right, or a request?
Open a discussion — happy to help.