gemma-4 12B coder for local, agentic tool use — GGUF quantizations for llama.cpp / Ollama.
Run it:llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja (full commands below).
⚠️ Tool-calling needs the recovery shim. The model emits gemma-4's native tool markup, which llama.cpp --jinja under-parses — wrap your endpoint with the tool-shim (see Tool-calling below) to get standard tool_calls.
💡 Pick this for the best tool-calling (our gate winner). For an uncensored model, use SFT v5 + abliterated GGUF.
1# llama.cpp (server) — tool-calling needs the recovery shim, see below2llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja --ctx-size 1638434# Ollama5ollama run hf.co/tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M
Files
Sizes and a one-click loader are in the file browser / Quantizations widget above;
the note says which quant to reach for.
Quant
Notes
Q4_K_M
good default — fits 12 GB VRAM, best size/quality balance
Tool-calling
Tool-calling works — but llama.cpp --jinja doesn't recognise gemma-4's native
tool-call markup, so the bare parser under-reports calls. The model is fine; the
parser is blind to the format. Recover standard tool_calls with a small serve-side
post-processor (no weight change, no latency beyond a regex scan).
Ready-to-use → tpls/gemma4-tool-shim — a drop-in
callback for OpenAI-compatible proxies, a standalone (dependency-free) example, and the pure
parser, all Apache-2.0, with the full recovery algorithm documented. Point your
OpenAI-compatible endpoint through it.
You send tools the usual OpenAI way (tools=[…]); the model emits native markup; the
shim turns it into a standard tool_calls object:
# model completion (raw):
<|tool_call>get_weather{"city": "Paris", "units": "celsius"}
json
1// after the shim:2{"finish_reason":"tool_calls",3"message":{"role":"assistant","content":null,4"tool_calls":[{"id":"call_0","type":"function",5"function":{"name":"get_weather","arguments":"{\"city\": \"Paris\", \"units\": \"celsius\"}"}}]}}
Tool-calling gate
Served the GGUF on llama.cpp (llama-server --jinja), prompted 7 tool-use cases + 1 no-tool abstain, scored whether a structured tool call was emitted. raw = llama.cpp native parse; shim = same outputs re-parsed for gemma-4 native markup. Tools folded into the prompt at eval time, matching training.
The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.
Measured on
Pass rate
this model — raw (--jinja)
0.125
this model — shim (prod path)
1.000
Intended use & limitations
Built for code generation and agentic tool use; serve locally via llama.cpp /
Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or
fabricated — validate tool arguments before executing, and keep a human in the loop
for anything consequential.
Salesforce/xlam-function-calling-60k (tools_mode=mixed, tools_ratio=0.5 (schemas folded into ~half the prompts)) @ 26d14ebfe18b1f7b524bd39b404b50af5dc97866
Training hyperparameters
Knob
Value
method
QLoRA (4-bit NF4 base, bf16 compute)
lora_r / lora_alpha / dropout
32 / 32 / 0.05
target_modules
all-linear
objective
train on assistant turns only (responses-only masking)
optimizer
adamw_8bit
lr / scheduler / warmup
2e-4 / cosine / 0.03
epochs
1
seq_len
4096
effective_batch
16
chat_template
gemma (native turn boundaries 105/106)
Training environment
Exact pins the run trained against (the base arch needs a recent transformers).
Package
Version
torch
2.11.0
transformers
5.13.0.dev0 @ c21da1b (git pin)
peft
0.19.1
trl
1.6.0
datasets
5.0.0
bitsandbytes
0.49.2
accelerate
1.14.0
liger-kernel
0.8.0
attention
eager (no flash-attn)
Quantization environment
The GGUF bytes depend on the quantizer build, not just the weights — a different
llama.cpp release rounds tensors differently and can change the convert mapping. Pins
the toolchain these quants were produced with:
llama-imatrix over the calibration set (CPU forward pass)
quantize
llama-quantize --imatrix, token-embeddings + output tensor kept at f16
The image is the rolling :full tag, not a digest — for byte-exact reproduction pin
the image digest you build with. The imatrix-quant step above lists the calibration
set and the EOG patch this build applied.