ZERO Gemma 4 E4B OpenZero — Standalone Compact Agentic GGUF
COMPACT ZERO. FULL STANDALONE MODEL. ONE DOWNLOAD.
Local research, coding and autonomous workflows without adapter setup.
GGUF
Gemma4
CPU
OpenZero
Zero Gemma 4 E4B OpenZero is a standalone, merged GGUF model for compact
local AI deployments. The trained OpenZero weights are already fused into the
model. There is no separate LoRA adapter and no second base-model download.
Small enough to run locally. Sharp enough to be Zero.
1ollama run hf.co/shafire/Zero-Gemma4-E4B-OpenZero-GGUF
2ollama run hf.co/shafire/Zero-Qwen3-8B-OpenZero-GGUF:Q5_K_M
OpenZero 7.1 can select the installed aliases locally. ZeroThink can route through its OpenZero Local provider when connected to an OpenZero node. Alias availability depends on which quantizations are installed on the machine.
OpenZero in action
OpenZero 7.1 panel showing Ultra mode and a 16-agent setting
OpenZero 7.1 — Ultra autonomy with the 16-agent setting.
ZeroThink Studio OpenZero local model selector
ZeroThink Studio — the three current OpenZero local model choices.
OpenZero local Agent Zero response
Local Agent Zero response inside the OpenZero control panel.
OpenZero agentic workflow in the Super Panel
OpenZero agentic workflow and local privacy controls.
Download this model
File
Size
Best for
Zero-Gemma4-E4B-OpenZero-Q5_K_M-F16-Merged.gguf
5.46 GiB
Recommended one-file release for local coding, research and agents
Standalone model: yes
Separate adapter required: no
Separate base model required: no
llama.cpp compatible: yes
OpenZero compatible: yes
CPU generation tested: yes
The file retains the Q5_K_M base tensors while preserving the 66 trained
attention tensors in F16. That keeps the fine-tuned tensors at higher precision
without forcing a second lossy quantization pass across the whole model.
What Zero is built for
Compact agentic AI: planning, structured execution and verification
Coding: debugging, implementation guidance, review and test design
Research: careful synthesis, evidence checks and explicit uncertainty
Tool workflows: deciding what to inspect, change and verify next
Local privacy: CPU-friendly inference without a hosted model dependency
OpenZero: OpenAI-compatible local serving for autonomous systems
The training material is already represented in the merged weights. Users do
not need the dataset, the training archive or a LoRA adapter to run this model.
Practical notes
Start with an 8K context on a 16 GiB system and increase it only after
measuring available memory.
This release is text-only. A multimodal projector is not included or required.
Tool execution is provided by the surrounding OpenZero or agent runtime.
The 512-token fine-tuning window strengthened targeted behavior; it does not
redefine all long-context behavior inherited from the base.
Validate high-stakes outputs independently.
OpenZero
Zero Gemma is the compact member of the Zero model family: local-first,
agent-oriented and built to verify before it boasts.