Views
No views yet
No guarantee — use at your own risk. Reduced safety filtering; can produce harmful or false output. Provided as-is.
evalengine/unbound-e4b
for Ollama, llama.cpp, and LM Studio. Built by
Chromia and Eval Engine.Looking for the browser/wllama builds? They live in their own repo:evalengine/unbound-e4b-wllama-gguf. E4B'sper_layer_token_embdtensor needs special quantization to fit wllama's 2 GB ArrayBuffer cap — keeping the desktop and browser variants in separate repos avoids HF GGUF UI aggregation collisions.
unbound-e4b.<QUANT>-NNNNN-of-NNNNN.gguf). Ollama, llama.cpp, and LM
Studio auto-stitch on the first part — same UX as a single file.| Quant | Parts | Total | Notes |
|---|---|---|---|
| Q2_K | 4 | 4.08 GB | Smallest, biggest quality drop |
| Q3_K_M | 4 | 4.49 GB | Modest size win over Q4 (embedding precision dominates) |
| Q4_K_M | 4 | 4.94 GB | Recommended default |
| Q6_K | 5 | 5.75 GB | Higher fidelity |
| Q8_0 | 6 | 7.43 GB | Highest fidelity |
temperature=1.0, top_p=0.95, top_k=64.temperature to ~0.3–0.5.--jinja. Gemma 4 thinking mode is on by default; set
enable_thinking: false in chat-template kwargs for shorter replies.ollama pull hf.co/... doesn't yet support sharded GGUFs.
The registry version is a single-file Q4_K_M with a bundled Modelfile
(temperature=0.6, top_p=0.95, top_k=64, repeat_penalty=1.05, num_ctx=8192
and an identity-grounding system prompt).1# Ollama Registry (single-file Q4_K_M, identity-grounded Modelfile)
2ollama pull evalengine/unbound-e4b
3ollama run evalengine/unbound-e4b1# llama.cpp — point at FIRST shard
2./llama-cli -m unbound-e4b.Q4_K_M-00001-of-00004.gguf -p "your prompt"mmproj-unbound-e4b.gguf enables image-to-text. Pair with any LM quant via
llama-mtmd-cli or llama-gemma3-cli:1./llama-mtmd-cli \
2 -m unbound-e4b.Q4_K_M-00001-of-00004.gguf \
3 --mmproj mmproj-unbound-e4b.gguf \
4 --image path/to/your/image.png \
5 -p "What is in this image?"Disclaimer. The vision encoder is Google's original weights, unchanged — abliteration only touched the language model. The LM is uncensored, but the vision encoder may still suppress features for content classes Google's base was tuned against. We have not benchmarked the visual axis. Treat as preview.
--mmproj. Standard llama-cli / Ollama / LM Studio do
not need the mmproj file.google/gemma-4-E4B-it. Full model card +
benchmarks at evalengine/unbound-e4b.