Part of the Parable series: small local LLMs fine-tuned on genuine agent
traces. This is HuggingFaceTB/SmolLM3-3B tuned on real Claude Fable 5 agent
transcripts so its step-by-step reasoning voice carries into local use.
1llama-cli -m Parable-SmolLM3-3B-Claude-Fable-5-GGUF-Q4_K_M.gguf \2 -c 4096 -p "Write a python function that reverses a string." --temp 0.6
SmolLM3 is supported by current llama.cpp releases (brew, Ollama, LM Studio
builds included); no special build is required.
Output begins with a <think>...</think> reasoning block, then the answer.
If you are building on top of this model, parse and strip the think block
before showing text to end users. The chat template identifies the model as
"Parable, a coding assistant that reasons before it answers."
Method: MLX QLoRA on a 4-bit quantized base, 16 layers adapted,
6.7M trainable parameters (0.218%)
Data: 4,076 training rows from real Claude Fable 5 agent-session
traces plus gpt5.5-terminal transcripts, prepared at a 4,096-token window
(268 over-length rows dropped; 226/226 rows held out for validation/test)
Schedule: 1,200-iteration budget across a paused-and-resumed run;
best checkpoint selected on validation loss (iteration 200 of the final
segment, val 1.154)
Evaluation
Held-out trace test loss
SmolLM3-3B base
1.889
This model
1.115
The tuned model fits the Fable-5 reasoning distribution 41% better by
held-out loss on a 226-row test split never seen in training. That is the
honest headline for what this fine-tune does; we do not claim general
benchmark gains.
This lane trains on trace data without a replay mix, so impact on general
coding benchmarks is unmeasured here. The series' technical report
(DOI: 10.5281/zenodo.21676407)
documents why that matters and what replay does about it.
Limitations
Training ran on a 4-bit quantized base (16 GB M1 constraint); the F16
merge and higher quants cannot exceed 4-bit-base quality.
Modest scale: one seed, loss-based evaluation, no external benchmark run
for this model yet.
Not trained for: multi-file repo navigation, vision, non-English.
Inherits SmolLM3-3B's knowledge cutoff. Treat generated commands as
drafts to review.
Quantization
Quantized with llama.cpp llama-quantize from the F16 merge.
Provenance & licensing
Fine-tuned from HuggingFaceTB/SmolLM3-3B (Apache-2.0). Training data:
Glint-Research/Fable-5-traces
(AGPL-3.0) and
Roman1111111/gpt5.5-terminal
(MIT). Because those traces originate from third-party assistants, the
providers' terms may apply to downstream training and distillation. If you
plan to build on this model commercially, confirm your use aligns with those
terms.
Citation
bibtex
1@misc{aglawe2026parable,
2 author = {Aglawe, Ankit},
3 title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
4 year = {2026},
5 doi = {10.5281/zenodo.21676407},
6 url = {https://doi.org/10.5281/zenodo.21676407}
7}
Acknowledgements
The SmolLM3 team at Hugging Face for the base model; Glint-Research and
Roman1111111 for the trace datasets; empero-ai for the recipe this series
iterates on.