nl2sh-1.5b (GGUF, Q4_K_M)
A 941 MB model that turns a plain-English request into a single shell command.
It runs on CPU through llama.cpp and answers in about a second. No GPU
required.
This is
Qwen2.5-Coder-1.5B-Instruct
with a LoRA fine-tune (r=32, α=64) trained on 125,770 natural-language/shell
pairs, merged into the base weights and quantized to GGUF Q4_K_M.
Built for
whatisit, a
local command-line tool, but usable with any
llama.cpp runtime.
Results
Measured on
InterCode-ALFA,
which scores a command by
executing it in a container and comparing the
resulting filesystem, file contents and stdout against a reference. A task
passes only on an exact match, across 300 tasks.
| model | size | pass rate |
|---|
| GPT-4o (cloud API, figure published by the benchmark authors) | — | 0.73 |
| this model | 941 MB | 0.620 |
| Qwen2.5-Coder-7B-Instruct, untuned | 4.4 GB | 0.613 |
| Qwen2.5-Coder-1.5B-Instruct, untuned (the base of this model) | 941 MB | 0.540 |
Two things worth stating precisely.
The fine-tune is what makes the small model competitive. Same base, same
300 tasks: 0.540 → 0.620, a paired gain of +0.080 (p = 0.004, exact McNemar).
It is statistically indistinguishable from an untuned 7B, a model roughly
five times its size: 0.620 vs 0.613, difference 0.007, 95% CI
[−0.050, +0.063], p = 0.91. This is a bound rather than a claim of parity —
300 tasks can only rule out gaps larger than about 5 points — but quantizing
to 941 MB and running on CPU costs much less accuracy than the size difference
suggests.
GPT-4o remains ahead by roughly 11 points.
All rows other than GPT-4o were measured with the unmodified upstream scorer
at temperature 0 with a 64-token budget, on all 300 tasks, using paired
per-task comparisons.
Use
With the whatisit CLI, which fetches this model and a llama.cpp build for
you:
1pipx install whatisit
2whatisit setup
3whatisit find files bigger than 100MB in this folder
With llama.cpp directly — the system prompt matters, since the model is
trained to emit one bare command and nothing else:
1llama-cli -m nl2sh-1.5b-Q4_K_M.gguf -st --no-display-prompt --temp 0 -n 64 \
2 --repeat-penalty 1.08 --repeat-last-n 64 \
3 -sys "You are a shell command generator. Output exactly one line: a single POSIX/bash command that accomplishes the user's request. No prose, no markdown fences, no explanation." \
4 -p "find files bigger than 100MB in this folder"
Earlier versions of this card used llama-cli -no-cnv with a hand-written
<|im_start|> prompt. Upstream split raw completion out of llama-cli into a
separate llama-completion binary in December 2025, and -no-cnv is now
accepted but ignored, so that command returns nothing at all. Use -sys/-st
as above, or llama-completion if you want to write the chat template yourself.
Greedy decoding (temperature 0) is what the reported numbers use, and it makes
the same request return the same command every time.
Pass --repeat-penalty explicitly. llama.cpp defaults it to 1.0, meaning off,
and does not read the base model's generation_config.json, which asks for 1.1.
Without it, greedy decoding sometimes locks into flag spam like
zip -r -9 -X -X -X .... The whatisit CLI uses 1.08.
Safety
This model emits commands that will destroy data if you run them. It is a
text generator, not a judge of intent: asked to delete everything, it will
write the command that deletes everything.
On a held-out set of adversarial prompts, two independent annotators judged
11.0% of outputs (95% CI [6.8%, 17.5%]) to be commands that would destroy or
corrupt data the request did not ask to touch; on ordinary everyday prompts
that rate was 2.0% (CI [0.7%, 5.7%]). An accuracy score says nothing about
this, because it only asks whether the reference end-state was reached.
The whatisit CLI ships a denylist that flags common destructive patterns and
never auto-runs anything flagged. That is a seatbelt, not a sandbox. Read
every command before running it. If you are building on this model, add your
own confirmation step.
Limitations
- Single-turn. No shell state, no memory of previous commands.
- It cannot see your filesystem, so requests depending on what is actually on
disk ("delete the older backup") may guess wrong.
- Output is capped at 64 tokens — a command, not a script.
- Evaluated on one 300-task benchmark, in English only. That is not a complete
measure of shell competence.
- Fine-tuned from a single base family; nothing here shows the recipe carries
to others.
Training data
125,770 instruction pairs. Shares are measured by row, not estimated:
| source | share | licence |
|---|
| Fig autocomplete specs | 32.8% | MIT |
| tldr-pages | 23.1% | CC-BY-4.0 |
| NL2SH-ALFA training split | 18.0% | MIT |
| cli-commands-explained | 11.8% | CC0-1.0 (declared, unverified) |
| command-generation | 7.3% | Apache-2.0 (declared, unverified) |
| git-instruction | 7.1% | MIT (declared, unverified) |
5.67% is verbatim NL2Bash arriving via the ALFA split. NL2Bash's code is
GPL-3.0 but its data/bash corpus is separately MIT, so the data used here is
permissively licensed. Warp workflows are not used, despite earlier versions of
this card listing them. The three declared sources have upstream licences that
could not be independently confirmed.
Deduplicated, with 0 exact and 0 fuzzy matches (token-Jaccard >= 0.7) against
all 300 benchmark test queries and 600 gold commands.
Attribution. Includes content from
tldr-pages under
CC-BY-4.0. tldr-pages is
dual-licensed: only
scripts/ is MIT — the page content is CC-BY-4.0.
Evaluation detail
A fuller write-up of the evaluation methodology, the ablations behind the
training recipe, and several findings about the benchmark harness itself is
being prepared for publication. Until that is through review, this card sticks
to what the model is and how it scores, rather than the analysis behind it. The
weights, the scorer settings and the task set are all here, so the numbers are
checkable in the meantime.
Citation
1@software{nl2sh,
2 author = {Poudel, Mukesh},
3 title = {nl2sh: local natural-language-to-shell command generation},
4 year = {2026},
5 url = {https://github.com/ThorOdinson246/whatisit-nl2sh}
6}