A small, on-device model is fast and private, but sometimes wrong. At Cactus we
post-train models to know when they are wrong: we ship probes inside the
checkpoint that score every answer with a confidence between 0 and 1,
returned as structured data (never parsed out of the answer text). Answer
on-device when confidence is high; re-route to a bigger model when it's low:
Gemma 4 E2B Hybrid, the smallest Gemma model, matches Gemini 3.1 Flash-Lite on
most benchmarks by routing only 15–35% of queries to Flash-Lite and running the
rest itself:
Benchmark
Handoff to match Flash-Lite (FP16)
At 4-bit
At 3-bit
ChartQA
15–20%
25–30%
40–50%
MMBench
30–35%
40–45%
50–55%
LibriSpeech
25–30%
35–40%
55–65%
GigaSpeech
30–35%
40–45%
50–55%
MMAU
30–35%
35–40%
50–55%
MMLU-Pro
45–55%
~90%
n/a
Quantisation quality is measured on
Cactus Quants,
which performs well at uniform quantization; developers are encouraged to
benchmark Unsloth, GGUF, and MLX quantization independently.
Quickstart
The gemma-4-e2b-it-hybrid architecture is not yet in upstream llama.cpp. Run
these files with a build that includes the Cactus patch series — on unpatched
llama.cpp they fail to load with "unknown model architecture" by design. Build
the patched server once:
bash
1git clone https://github.com/cactus-compute/cactus-hybrid &&cd cactus-hybrid
2./patches/llama.cpp/install.sh && rehash # clones the pinned tag, applies the patches, builds
1curl -s http://localhost:8080/v1/chat/completions \2 -d '{"messages":[{"role":"user","content":"What is the capital of France?"}],"max_tokens":512}'\3| jq '{answer: .choices[0].message.content, confidence}'
Chat-completions responses (and the final SSE chunk when streaming) carry a
top-level "confidence" field.
Files
file
quant
size
notes
gemma-4-e2b-it-hybrid-f16.gguf
F16
9.31 GB
closest to the bf16 reference
gemma-4-e2b-it-hybrid-Q4_K_M.gguf
Q4_K_M
3.43 GB
recommended for consumer hardware
The probe head (11 probe.* tensors) is stored in F32 in all quants —
only the trunk is quantized.
Calibration note
Quantized trunks shift the layer-28 activations the probe reads, moving
confidences downward relative to the bf16 reference (measured mean drift:
F16 −0.07, Q4_K_M −0.10; easy-vs-hard ordering fully preserved). If you use
aggressive thresholds, calibrate per quant; the 0.85 default remains
conservative (it hands off more, never less).
Routing quality (AUROC)
AUROC measures how well the probe separates wrong answers from right ones
(higher = better, 0.5 is random, 1.0 is perfect):
Hold-out
Modality
Cactus Hybrid
Token Entropy
MMLU
text MCQ
0.770
0.697
MMLU-Pro
text MCQ
0.771
0.692
ARC-Easy
text MCQ
0.888
0.655
ARC-Challenge
text MCQ
0.834
0.646
GSM8K (3-shot)
text gen
0.782
0.731
MMBench-EN-Dev
vision MCQ
0.840
0.435
ChartQA
vision QA
0.779
0.615
DocVQA
vision QA
0.781
0.512
MMAU
audio MCQ
0.789
0.517
GigaSpeech
audio
0.876
0.343
Earnings-22
audio
0.839
0.323
LibriSpeech
audio
0.822
0.427
Mean
0.814
0.549
The strongest result: the probe was trained on zero audio data, yet achieves
0.79–0.88 AUROC on four audio benchmarks (two transcription, one audio MCQ, one
out-of-domain transcription). This rules out surface-level explanations: the
probe is reading a modality-independent correctness signal from the hidden
state, not memorizing patterns from training data.