This model is the result of an effort to improve a small coding model by reducing the token count for reasoning traces while retaining accuracy. It worked surprisingly well - 74% shorter reasoning tokens on math prompts. Smaller reductions to actual code, but accuracy also improved.
This is a reproduction of the ThinkingCap method (BottleCap AI) by doing GRPO training with correctness reinforcement and length penalized.
I ran this tune using Claude Code driving the training based on a question - can we apply ThinkingCap to coding models and improve their performance?
TL;DR
What: RL fine-tune that cuts reasoning-token spend while holding correctness.
How much: −29–30% total tokens across GSM8K/HumanEval/MBPP; −74% on GSM8K specifically.
Accuracy: held or improved everywhere (avg 83.9% → 85.5%). No regression.
Behavior: does not loop or degenerate — it seems to be calmer than the base model under
ambiguity (see loop-rate eval).
Base: Ornith-1.0-9B (MIT) → post-trained on Qwen3.5 (Apache 2.0). This model: MIT.
See License for the full chain.
Usage
This is a LoRA adapter (222 MiB), not a full model — apply it on top of the base
deepreinforce-ai/Ornith-1.0-9B. Want a ready-to-run quantized build instead? See the GGUF
companion repo khudgins/Ornith-1.0-9B-ThinkingCap-GGUF (Q8_0 recommended).
Quick use — transformers + PEFT:
python
1import torch
2from transformers import AutoModelForCausalLM, AutoTokenizer
3from peft import PeftModel
45BASE ="deepreinforce-ai/Ornith-1.0-9B"6ADAPTER ="khudgins/Ornith-1.0-9B-ThinkingCap"# this repo78tok = AutoTokenizer.from_pretrained(BASE)9model = AutoModelForCausalLM.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")10model = PeftModel.from_pretrained(model, ADAPTER).eval()1112msgs =[{"role":"user","content":"A store had 48 apples, sold 3/4, then got 30 more. How many now?"}]13inp = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",14 return_dict=True).to(model.device)15out = model.generate(**inp, max_new_tokens=1024, do_sample=False)16print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=True))
Hosting — vLLM (LoRA served on the base, hot-swappable):
bash
1vllm serve deepreinforce-ai/Ornith-1.0-9B \2 --enable-lora \3 --lora-modules thinking-cap=khudgins/Ornith-1.0-9B-ThinkingCap \4 --max-lora-rank 325# then call the OpenAI-compatible API with "model": "thinking-cap"
1# Ollama — pulls the GGUF straight from HF2ollama run hf.co/khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 "your prompt"34# llama.cpp — run, or serve an OpenAI-compatible endpoint5llama-cli -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0 -p "your prompt"6llama-server -hf khudgins/Ornith-1.0-9B-ThinkingCap-GGUF:Q8_0
For a permanent single-model HF/vLLM deployment you can instead merge the adapter into the
base (PeftModel.merge_and_unload()) — but serving the adapter keeps the base swappable and the
download tiny. Building your own GGUF from a merged model? Set mtp_num_hidden_layers=0 in its
config.json first, or the converter counts a phantom 33rd block and the GGUF won't load.
Results
Greedy decode, evaluated in-process (transformers + PEFT). Each cell is
accuracy @ mean completion tokens.
Benchmark (N)
Base Ornith-9B
Thinking-Cap (this model)
GSM8K (150)
88.0% @ 533
91.3% @ 136 (−74% tokens)
HumanEval (164)
92.7% @ 728
92.1% @ 669 (−8% tokens)
MBPP (100)
71.0% @ 846
73.0% @ 689 (−19% tokens)
average accuracy
83.9%
85.5%
total tokens
2107
1494 (−29%)
Adaptive concision. The token cut scales to how compressible each task is — dramatic on
math (reasoning is cheap), modest on code.
Does it loop? No.
A frequent failure mode of length-penalized RL is repetitive degeneration ("anxiety") in
agentic use. Tested on 24 loop-prone prompts (hard reasoning, eval-framing, ambiguous),
greedy decode, with a repeated-span detector:
Model
Loop rate
Mean tokens
Runaway completions
Base Ornith-9B
0.0%
1014
3
Thinking-Cap
0.0%
607
1
Neither loops; the tune produces fewer runaway completions than the base model. The correctness
reinforcement kept the accuracy up.
Quantization
GGUF quants preserve the behavior. Benchmarked via Ollama (one harness for both, so the
difference isolates the quantization effect; absolute numbers run slightly different from the
transformers table above because of the GGUF chat-template path):
Format
Size
GSM8K
HumanEval
MBPP
avg
f16
17 GB
91.3%
90.9%
77.0%
86.4%
Q8_0
9.5 GB
90.0%
89.0%
76.0%
85.0%
Q8 lands within 1–3 problems of f16 on every benchmark (−1.4 pts average) with identical token
behavior — quantization is effectively free. Q8_0 at half the memory is the recommended
daily driver; f16 is the publish-precision reference. (Lower quants are not recommended: a Q4
of the base model failed to complete the eval at all.)
Method
Correctness-gated, length-penalized GRPO (RL with verifiable rewards). Per completion:
reward = correctness # 1.0 if the answer/tests pass, else 0.0
− (λ · normalized_length if correct) # length penalty ONLY when correct
+ format_bonus # well-formed <think>…</think> + answer
The gate is the safety property. The length penalty applies only to correct
answers, so the model can never trade accuracy for brevity.
GRPO's group-relative advantage means compression only happens on prompts the model has
already mastered; a KL leash to the frozen base model and a λ ramp (off → max over
~300 steps) keep it stable and prevent degeneration. The result is self-limiting: the length
curve hits a floor and plateaus (checkpoints at steps 400–800 are a tied plateau).
Trained with LoRA (r32/a64, all-linear, MoE router frozen) via Unsloth + TRL GRPOTrainer
on an NVIDIA DGX Spark (GB10). Reward signal reuses the same verifiable checkers used to grade
the model, so training and evaluation share one oracle.
Intended use & limitations
Use: general reasoning and coding assistant where you want the model's accuracy without
paying for verbose chain-of-thought — interactive coding (works well in agent frontends),
math, tool use.
Limitations: English-focused, inherits base Ornith/Qwen3.5 limitations. The compression
is strongest on math-style reasoning; code responses stay appropriately detailed. As with any
9B model, verify security- or correctness-critical output.
Not speculative-decoding capable in this GGUF (MTP head dropped for llama.cpp compatibility).
Reproduction & provenance
Full recipe, reward harness, and eval scripts: github.com/khudgins/ornith-thinking-cap. Method:
ThinkingCap (correctness-gated length-penalized GRPO) — an independent reproduction of the
ThinkingCap model series by BottleCap AI
(announcement ·
bottlecapai/ThinkingCap-Qwen3.6-27B),
applied here to a smaller coding model. Not affiliated with or endorsed by BottleCap AI. Base model:
deepreinforce-ai/Ornith-1.0-9B (MIT per its HF tag), itself post-trained on Qwen3.5 (Apache
2.0). See License for the full chain. Tooling: Unsloth, TRL, llama.cpp, Ollama.
License
MIT. This fine-tune — the LoRA adapter and any weights merged from it — is released under
the MIT License. It adds no new base weights; it's a low-rank delta trained on top of the base.
Upstream provenance and terms, stated precisely:
Base model:deepreinforce-ai/Ornith-1.0-9B
is tagged MIT on Hugging Face; the inherited base weights are governed by that grant.
Underlying architecture: Ornith-1.0-9B is post-trained on Qwen3.5, released under
Apache 2.0. Apache 2.0's terms (notably its patent grant and NOTICE requirements) apply to
the Qwen3.5-derived weights inherited through the base.
Not a Gemma derivative. The Ornith family card also lists Gemma-4-based variants, which
would carry Google's Gemma license. This model is not one of those — it derives solely from
the Qwen3.5-based, MIT-tagged Ornith-1.0-9B (verified: Qwen3_5ForConditionalGeneration).