Tool-Use Benchmark for Small/Efficient LLMs on Apple Silicon
Benchmarking tool-use (function calling) capability of small and efficient language models running locally on a Mac Mini M4 (16 GB). Covers 13 model configurations across 3 backends, evaluated on 3 benchmarks.
Key Finding
PrismML Bonsai-8B (1-bit, 1.15 GB) achieves 73.3% on BFCL — beating Qwen3.5-9B (61%), Gemma 4 E4B (65%), and every other model tested. A 1-bit quantized model at 14x compression outperforms 4-bit models 4-5x its size on structured function calling.
However, on complex API semantics (NexusRaven), Bonsai-8B drops to 43.8% vs Qwen3.5-9B's 75%. BFCL measures format compliance; NexusRaven measures understanding.
Results
BFCL (Berkeley Function Calling Leaderboard) — 50 tests per category
Model
Size
Quant
Backend
Simple
Multiple
Parallel
Avg
Avg Time
Bonsai-8B
1.15 GB
Q1_0 1-bit
llama.cpp
68%
72%
80%
73.3%
1.8s
Gemma 4 E4B-it
~5 GB
Q4_K_M
Ollama
54%
64%
78%
65.3%
2.4s
Qwen3.5-9B
~5 GB
Q4_K_M
llama.cpp
56%
68%
68%
64.0%
11.6s
Qwen3.5-9B
~5 GB
MLX 4-bit
mlx-vlm
60%
68%
64%
64.0%
9.5s
Qwen2.5-7B
~4.7 GB
Q4_K_M
Ollama
58%
62%
70%
63.3%
2.9s
Gemma 4 E2B-it
~3 GB
Q4_K_M
Ollama
56%
60%
70%
62.0%
1.3s
Gemma 3 12B
~7.3 GB
Q4_K_M
Ollama
54%
54%
78%
62.0%
5.4s
Qwen3.5-9B
~5 GB
Q4_K_M
Ollama
50%
60%
74%
61.3%
5.4s
Bonsai-4B
0.57 GB
Q1_0 1-bit
llama.cpp
36%
56%
74%
55.3%
1.0s
Bonsai-1.7B
0.25 GB
Q1_0 1-bit
llama.cpp
58%
54%
54%
55.3%
0.4s
Llama 3.1 8B
~4.7 GB
Q4_K_M
Ollama
46%
42%
66%
51.3%
3.0s
Mistral-Nemo 12B
~7.1 GB
Q4_K_M
Ollama
40%
44%
64%
49.3%
4.4s
Bonsai-4B FP16
7.5 GB
FP16
mlx-lm
8%
34%
34%
25.3%
4.8s
NexusRaven API Evaluation — 48 queries, 12 per domain, stratified
Model
Size
Overall
cve_cpe
emailrep
virustotal
toolalpaca
Avg Time
Qwen3.5-9B (llama.cpp)
~5 GB
77.1%
58%
100%
100%
50%
14.1s
Qwen3.5-9B (Ollama)
~5 GB
75.0%
58%
100%
100%
42%
4.1s
Qwen2.5-7B
~4.7 GB
70.8%
50%
92%
100%
42%
2.0s
Qwen3.5-9B (mlx-vlm)
~5 GB
70.8%
50%
100%
92%
42%
13.8s
Gemma 3 12B
~7.3 GB
68.8%
33%
100%
100%
42%
3.5s
Llama 3.1 8B
~4.7 GB
66.7%
25%
92%
100%
50%
2.1s
Mistral-Nemo 12B
~7.1 GB
66.7%
42%
92%
100%
33%
3.0s
Gemma 4 E4B-it
~5 GB
60.4%
33%
83%
83%
42%
1.6s
Bonsai-1.7B
0.25 GB
54.2%
25%
75%
83%
33%
0.3s
Gemma 4 E2B-it
~3 GB
47.9%
17%
67%
75%
33%
0.9s
Bonsai-4B
0.57 GB
43.8%
17%
58%
75%
25%
0.8s
Bonsai-8B
1.15 GB
43.8%
17%
67%
67%
25%
1.2s
Bonsai-4B FP16
7.5 GB
29.2%
8%
42%
50%
17%
3.5s
AgentBench OS — Multi-step agentic tasks in Docker containers
Model
Score
Backend
Qwen3.5-9B
4.5/10 (45%)
Ollama / llama.cpp (tied)
Qwen2.5-7B
3.5/10 (35%)
Ollama
Mistral-Nemo 12B
3.5/10 (35%)
Ollama
Llama 3.1 8B
3.0/10 (30%)
Ollama
Gemma 3 12B
2.5/10 (25%)
Ollama
Backend Comparison (Qwen3.5-9B only, all 3 benchmarks)
Backend
AgentBench
BFCL
NexusRaven
Composite
llama.cpp (UD-Q4_K_XL)
45%
64.0%
77.1%
62.0%
Ollama (Q4_K_M)
45%
61.3%
75.0%
60.4%
mlx-vlm (MLX-4bit)
42%
64.0%
70.8%
58.9%
Hardware
Mac Mini M4 (10-core CPU, 10-core GPU, 16 GB unified memory)
macOS Sequoia 15.3
All models run locally, no cloud APIs
Benchmarks
1. BFCL (Berkeley Function Calling Leaderboard)
What: Structured function calling — can the model output the correct function name and parameters?
tool-use-bench/
README.md
scripts/
run_bfcl.py # BFCL evaluator (multi-backend)
run_nexusraven.py # NexusRaven evaluator (multi-backend)
run_bench.py # AgentBench OS evaluator (Docker)
results/
bfcl/ # 13 result JSONs (per model+backend)
nexusraven/ # 13 result JSONs
agentbench/ # Per-backend and per-model results
ollama/
llama-cpp/
mlx-vlm/
qwen2.5_7b/
gemma3_12b/
llama3.1_8b/
mistral-nemo_12b/
datasets/
agentbench_os_v1.json # 10 OS interaction tasks (set 1)
agentbench_os_v2.json # 10 OS interaction tasks (set 2)
agentbench_os_v3.json # 10 OS interaction tasks (set 3)
agentbench_os_v4.json # 10 OS interaction tasks (set 4)
Insights
1-bit quantization can preserve tool-use capability — Bonsai's quantization-aware training actually improves structured output (73% BFCL) compared to its FP16 base (25%). The 1-bit model is both smaller and better.
BFCL and NexusRaven measure different things — Bonsai excels at format compliance (BFCL) but struggles with complex API semantics (NexusRaven). For edge deployment where API schemas are fixed and simple, Bonsai is excellent. For dynamic/complex APIs, Qwen3.5-9B is the better choice.
Backend choice barely matters — Ollama, llama.cpp, and mlx-vlm produce nearly identical accuracy (within ~3%). The model is the bottleneck, not the serving infrastructure.
Size vs capability trade-off — Bonsai-1.7B (0.25 GB, 0.4s/query) scores 55% on BFCL and 54% on NexusRaven. For the price of a single image, you get a competent function-calling model that runs on a phone.
License
Scripts are MIT licensed. Benchmark datasets are subject to their original licenses: