Views
No views yet
Note on BFCL scores: All BFCL scores reported here use our internal simplified evaluation (single-function-call subset with custom prompt/scoring), NOT the official Berkeley Function Calling Leaderboard methodology. Our scores are not directly comparable to the official leaderboard. We are working on running the official BFCL evaluation for comparable numbers.
| Model | Size | Active | HumanEval+ | MBPP+ | BFCL v3 | NL2Bash F1 |
|---|---|---|---|---|---|---|
| Gemma 4 31B IQ2_M | 10.4 GB | 31B | 88.41% | 82.01% | 92.25% | 84.71% |
| This model (26B MoE IQ2_M) | 9.7 GB | 3.8B | 80.49% | 77.78% | 92.75% | 79.64% |
| Qwen3.6 IQ2_M | 11.1 GB | 3B | 80.49% | 78.31% | 94.75% | 81.63% |
| Gemma 4 E4B Q8_0 | 7.8 GB | 4.5B | 73.78% | 73.28% | 93.75% | 79.75% |
1huggingface-cli download KikoCis/gemma-4-26B-A4B-it-IQ2_M-GGUF gemma4-26b-a4b-IQ2_M.gguf --local-dir .
2
3# Without thinking
4llama-cli -m gemma4-26b-a4b-IQ2_M.gguf -ngl 99 --ctx-size 8192 --temp 0.1 \
5 -p "Write a Python function to merge two sorted lists"
6
7# With thinking (recommended for complex tasks)
8llama-cli -m gemma4-26b-a4b-IQ2_M.gguf -cnv -ngl 99 --ctx-size 8192 \
9 --reasoning on --reasoning-budget 1024gemma4-26b-a4b-IQ2_M.gguf — quantized weights (9.7 GB)gemma4-26b-a4b-domain.imatrix — importance matrixBenchmark scores do not predict agent capability. In Docker-based autonomous testing, fine-tuned E4B models (95% BFCL) scored 0/10 while the unfine-tuned base scored 6/10. Fine-tuning for BFCL destroyed general reasoning (error recovery, strategy adaptation, anti-repetition). Fine-tuned E4B models have been withdrawn.For autonomous agent tasks, use the base Gemma 4 model or a larger model at higher BPW. See: The Benchmark Trap — Full Study