Benchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 72.2% with Python tool.
Metric
Value
Correct
722
Incorrect
261
Errors
17
Total samples
1000
Python tool calls
915
Python tool errors
0
Total completion tokens
872,102