Evaluation of a Qwen3-1.7B-Base SFT checkpoint trained on 10K SWE-ZERO trajectories, on the 100-task SWE-bench Verified slice from marin#4898.
pass@1 = 7/100 = 7% (latest trial per task)
For reference: pass@N (any of N attempts won) = 9/100 = 9%. Some tasks received multiple trial attempts due to TPU preemption; only the latest complete attempt counts for pass@1.
pass@1 resolved task (latest trial… See the full description on the dataset page:
https://huggingface.co/datasets/AlienKevin/SWE-ZERO-10K-Qwen3-1.7B-Base-eval.