A 4B local coding model with agent instincts. Planning, tool habits and
terminal reasoning distilled from real Claude Fable 5 agent sessions, not
synthetic Q&A. Runs on ~2.5 GB of RAM.
ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M
v3.1 (2026-08-11)
Retrained on corpus v3.1: the v2 agent traces plus 1,807 execution-verified
solutions generated by the previous build and kept only where the code actually
ran against its tests. Two seeds souped, merged at the v2.1 scale.
The result matches or beats base Qwen3-4B on all four execution benchmarks,
where the previous build trailed it on three. If you pulled this model before
11 August 2026, re-pull.
Files
File
Quant
Size
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.gguf
Q4_K_M
2.5 GB
recommended
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q5_K_M.gguf
Q5_K_M
2.9 GB
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q6_K.gguf
Q6_K
3.3 GB
Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q8_0.gguf
Q8_0
4.3 GB
Parable-Qwen3-4B-Claude-Fable-5-GGUF-F16.gguf
F16
8.1 GB
for re-quantizing
What it is good at
It answers. Base Qwen3-4B spends its whole budget inside <think> on
34% of ordinary prompts and returns nothing. This model answers 34/34 on
the same suite, with 140x less reasoning text and no thinking-mode flag to
manage.
Agent-shaped reasoning. Trained on genuine multi-step agent sessions,
so plans, tool selection and terminal workflows come out structured
instead of improvised.
Small enough to keep open. Q4_K_M is 2.5 GB. Laptop, old GPU, modest
desktop — it runs, offline, with your code staying on your machine.
Evaluation
Measured on identical harnesses, greedy decoding, Q4_K_M builds, thinking
disabled on every row. Base and this model run through the same instrument in
the same session.
Base Qwen3-4B
This model (v3.1)
HumanEval
73.2
74.4
HumanEval+
68.3
68.3
MBPP
69.0
72.8
MBPP+
59.8
63.8
Held-out agent-trace loss
2.155
1.446
Measured on the v2.1 build and carried forward (the training objective and
chat behaviour are unchanged):
Base Qwen3-4B
Parable
Prompts answered (34-prompt suite)
27/34
34/34
BFCL simple_python
95.3
92.3
BFCL multiple
94.5
90.0
Choosing between this and the base
Take this model for local agent and coding work where you want
structured, reliable answers every time: it fits the agent-session
distribution far better and never silently returns empty.
Take the base model if your workload is maximum-accuracy function
calling in a tool-calling harness, where its few extra points matter more
than reasoning style.
Fine-tuned from Qwen/Qwen3-4B (Apache-2.0). Training data:
Glint-Research/Fable-5-traces
(AGPL-3.0) and
Roman1111111/gpt5.5-terminal
(MIT). Because those traces originate from third-party assistants, the
providers' terms may apply to downstream training and distillation. If you
plan to build on this model commercially, confirm your use aligns with those
terms.
Citation
bibtex
1@misc{aglawe2026agenttrace,
2 author = {Aglawe, Ankit},
3 title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
4 year = {2026},
5 publisher = {Zenodo},
6 doi = {10.5281/zenodo.21676407},
7 url = {https://doi.org/10.5281/zenodo.21676407}
8}
Acknowledgements
The Qwen team for the base model; Glint-Research and Roman1111111 for the
trace datasets; empero-ai for the recipe this series iterates on.