How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page:
https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.