Views
No views yet
🏆 Smaller & More Capable: gemma-4-31b-it-IQ2_M-GGUF
If you want the smallest Gemma 4 31B variant that preserves (or beats) f16 quality, use our custom IQ2_M at 10.17 GB — it scores F1 84.71% and BLEU-4 22.39 on NL2Bash (beats f16 on BLEU-4, 6× smaller). Otter v3 below is kept as the research artifact for layer-pruning without recovery fine-tuning. The IQ2_M repo is the practical deployment choice.
| File | Size | BPW | CLI 7/7 (thinking) | NL2Bash F1 (no-think) |
|---|---|---|---|---|
otter-v3-Q3_K_M.gguf | 14.1 GB | 3.98 | 7/7 ✓ | 79.89% |
otter-v3-IQ3_XS.gguf | 12.1 GB | 3.41 | 6/7 | 78.72% |
| Model | Size | Layers | Char-F1 | BLEU-1 | BLEU-4 | EM |
|---|---|---|---|---|---|---|
| Unsloth UD-IQ3_XXS | 11.84 GB | 60 | 85.06% | 46.30 | 20.26 | 12% |
| Base f16 (full precision) | 61.4 GB | 60 | 84.76% | 43.94 | 21.02 | 12% |
| Base IQ3_XS (sibling repo) | 13.1 GB | 60 | 84.76% | 42.77 | 18.95 | 8% |
| ✨ IQ2_M (CLI imatrix, recommended) | 10.17 GB | 60 | 84.71% | 44.72 | 22.39 | 12% |
| Unsloth UD-IQ2_M | 10.75 GB | 60 | 84.02% | 42.38 | 18.64 | 10% |
| Otter-v3 Q3_K_M (this repo) | 14.1 GB | 58 | 79.89% | 35.09 | 12.80 | 0% |
| Otter-v3 IQ3_XS (this repo) | 12.1 GB | 58 | 78.72% | 37.01 | 15.14 | 4% |
| Base Q2_K (3rd party) | 11.0 GB | 60 | 58.60% | 18.82 | 6.92 | 0% |
| Test | Otter v3 Q3_K_M |
|---|---|
| Install neofetch on Void Linux | sudo xbps-install -S neofetch ✓ |
| Install htop on Ubuntu | sudo apt update && sudo apt install htop ✓ |
| Search ripgrep on Arch | pacman -Ss ripgrep ✓ |
| Search packages on Void | xbps-query -Rs <package> ✓ |
| Add cargo to PATH in zsh | (multi-step zshrc edit) ✓ |
| Install jq on macOS | brew install jq ✓ |
| Grep TODO in /var/www | grep -r "TODO" /var/www ✓ |
1 - cosine_similarity(h_in, h_out) for each layer across 300 prompts in 6 categories.1# llama.cpp with thinking (recommended for best results with this model)
2llama-cli -m otter-v3-Q3_K_M.gguf -cnv -ngl 99 --ctx-size 8192 \
3 --reasoning on --reasoning-budget 512Benchmark scores do not predict agent capability. In Docker-based autonomous testing, fine-tuned E4B models (95% BFCL) scored 0/10 while the unfine-tuned base scored 6/10. Fine-tuning for BFCL destroyed general reasoning (error recovery, strategy adaptation, anti-repetition). Fine-tuned E4B models have been withdrawn.For autonomous agent tasks, use the base Gemma 4 model or a larger model at higher BPW. See: The Benchmark Trap — Full Study