grug put 27b brain in four-bit cave DURING training. brain feel rounding rock
before final squish. this not normal quant. this QAT recovery rock, made for
Q4 people.
recipe (same as grug-9b-qat, scaled): full-weight QAT on
grug-27b, fake int4 asymmetric
group-32 with straight-through gradient, ~3M token of grug-think data,
Adafactor LR 2e-6, 249 step on one H200. release = 25% QAT move + 75% original
anchor (full QAT overcorrect, 9b teach grug this). then fresh BF16 export,
quantize ONE time to Q4_K_M.
three rock fight: ordinary Q4 (control), full-QAT Q4, and this rock (25%
QAT blend). full-QAT win MBPP big but BREAK tool hand (agent valid 96.6 ->
82.8). grug no ship broken hand. blend rock best overall:
QAT feel rounding rock during training -> Q4 squish hurt less. coding and math
UP, loop sickness DOWN, tool hand intact. small right-tool dip is the honest
trade. grug show all numbers, hide nothing.
single-token spam ("/" forever etc) = NOT the rock. hybrid DeltaNet brain
CANNOT survive llama.cpp context-shift: old builds shift on context overflow
and corrupt the recurrent state into token spam. fix:
1llama-server -m grug-27b-qat-Q4_K_M.gguf --mmproj mmproj-grug-27b-f16.gguf \
2 -c 16384 --temp 0.6 --top-p 0.95 --top-k 20
need recent llama.cpp (qwen3_5 arch). grug think live in
<think>, short on
purpose. tool call use XML format. main model card:
grug-27b.