old grug-35b repeat words until cave fall down. grug bury him, study bones,
rebuild WHOLE brother from zero with the grug-27b pipeline. v2 different
animal: complete retrain, repetition gates, stripped-history training, deep
grug reasoning. then v2.1 round on top (long-hunt + deep-think data).
base: Ornith-1.0-35B
(mixture-of-experts) - template keep think in every turn, grug respect that
method: LoRA on attention + DeltaNet + shared paths (expert FFN and
router untouched - grug no poke router), merged bf16, then v2.1 continued
round with long agent sessions + deep-think data
grug think in dense grug-speak inside <think>, answer normal english
number. same harness as grug-27b, all open
test
Ornith-35B base
grug-35b-v2.1
token (base -> grug)
HumanEval
86.6
84.8
4882 -> 215 (-96%)
MBPP
92.0
88.0
1883 -> 96 (-95%)
GSM8K
92.5
92.0
1512 -> 105 (-93%)
MATH-500 (unseen, base give 12k budget)
63.3
64.7
6892 -> 209 (-97%)
SWE replay: right tool %
52.9
79.4
SWE replay: args schema-valid %
100
100
repetition gauntlet healthy
83.7
86.0
degenerate loops
0
0
think blocks closed
93.0%
100%
MATH-500 special: never touch any grug data pipeline. base got 3x bigger
budget and still 32% of base run hit token wall. grug beat base with 33x
fewer token. no benchmaxx, just no fat.
grug honest corner: HumanEval -1.8, MBPP -4 vs base. and tool-call PRESENCE
82.4% vs base 98.5 (grug sometimes answer in words when tool call better;
when grug DO call, args 100% valid and right tool way more often, 79.4 vs
52.9). old v1 repetition sickness: extinct. zero loops all gauntlets.
v2 -> v2.1 changelog
test
v2
v2.1
MBPP
86.0
88.0
right tool %
77.9
79.4
GSM8K
93.5
92.0
loops
0
0
v2.1 add: stuck-loop escape, sacred no-tool final summary, fresh think every
turn, deep-think on hard problems. same medicine as grug-27b v2.1.