Run ID: 20260704-swe_bench_verified_mini-openclaw-gemma-4-31b-it-fixedfull2
Benchmark: SWE-Bench Verified Mini
Agent/harness: openclaw via AlphaDiana Podman SWE harness
Model: google/gemma-4-31B-it
Slurm job: fixedfull2
Summary from local inspection:
50 task rows
49 valid_scored
1 runtime_error
0 provider_error
17 correct
finish reasons: timeout=8
valid-only accuracy: 0.3469
completed-row accuracy: 0.3400
total-denominator… See the full description on the dataset page:
https://huggingface.co/datasets/n-pelleriti/alphadiana-swe-bench-mini-completed-results.