This dataset is a small, curated subset of a larger internal benchmark designed for training and evaluating tool routing behavior in lightweight language models.
It is used in the development of the Kumru model family, specifically for improving the model’s ability to decide:
when to call a tool
when to respond directly
Metrics are after applying fine-tuned adapter, base model remains unchanged.
image
Model
Num Examples
Valid JSON Rate
Schema Valid Rate
Route Accuracy
Kumru Base
218
0.8578
0.8532
0.4908
Kara Kumru
218
0.9450
0.9128
0.4908
Kumru Kuymak v1.1
218
1.0000
1.0000
0.8532
Model
Valid JSON Δ
Schema Valid Δ
Route Accuracy Δ
Kumru Base
base
base
base
Kara Kumru
+10.16%
+6.98%
+0.00%
Kumru Kuymak v1.1
+16.57%
+17.20%
+73.83%
Dataset Overview
Language: Turkish
Task type: Binary classification (Tool vs Direct response)