Views
No views yet
google/gemma-4-31B-it. One point in a rank x LR sweep on the Rovo Chat Orchestrator tool-calling task.| metric | this model | base gemma-4-31B-it |
|---|---|---|
| Text AA (shallow) | 89.3% | 96.1% |
| Tool AA (shallow) | 53.2% | 53.2% |
| Overall AA | 79.9% | 84.8% |
| Text AA Deep (binary, 1-5 rubric) | 74.8% | 86.7% |
| Deep quality mean (1-5) | 4.09 | 4.27 |
| False-tool-trigger rate | 7.8% | 1.9% |
| Runaway rate | 1.4% | 0.3% |
tzchen07/gemma4-31b-rovochat-lora-r16-lr5e5).-it model overall — narrow-domain tuning erodes the base's tool-vs-text calibration faster than it gains tool accuracy. Base wins on text AA, overall, and stability; the best tune improves only tool AA (+3 pts). Companion full-FT analysis (LR is the killing factor): tzchen07/gemma4-31b-rovochat-sft-v3.