Views
No views yet
<think> into the prompt itself:{{ '<|im_start|>assistant\n<think>' }}reasoning_content, and if
you apply a response_format schema the grammar takes the content stream from
the first token - which is the same stream the model wants to think in. It then
does not think at all.1-{{ '<|im_start|>assistant\n<think>' }}
2+{{ '<|im_start|>assistant\n' }}| Size | ||
|---|---|---|
olmo-3-7b-think-q4_k_m.gguf | 4.5 GB | the only one so far, ask if you want Q8_0 |
1llama-server -m olmo-3-7b-think-q4_k_m.gguf --ctx-size 32768 \
2 --reasoning on -ngl 99--max-tokens 20000 is not excessive, and a small budget gets
you an empty answer rather than a short one.gguf_new_metadata.py from llama.cpp b10223, rewriting only
tokenizer.chat_template on
lmstudio-community/Olmo-3-7B-Think-GGUF.
All 355 tensors are theirs, unchanged.