Run ID: 20260602-swe_bench_verified_mini-opencode-gemma-4-31b-it-full-818023
Benchmark: SWE-Bench Verified Mini
Agent/harness: OpenCode via AlphaDiana Podman SWE harness
Model: google/gemma-4-31B-it
Slurm job: 818023
Summary from local inspection:
50 task rows
35 valid_scored
14 provider_error context overflows
1 runtime_error image build failure
19 correct
valid-only accuracy: 19/35 = 0.5429
total-denominator pass@1/accuracy: 19/50… See the full description on the dataset page:
https://huggingface.co/datasets/n-pelleriti/alphadiana-swe-mini-opencode-gemma4-20260602-818023.