Benchmark of google/gemma-4-E2B-it against MMLU-Pro dataset. Model's answer is considered correct if it matches the ground truth answer index exactly.
Accuracy: 61.6% with Python tool.
Metric
Value
Correct
617
Incorrect
381
Errors
3
Total samples
1001
Python tool calls
314
Python tool errors
12
Total completion tokens
1,499,382