Evaluation of a Qwen3-1.7B-Base SFT checkpoint trained on 1M SWE-ZERO trajectories with full-transcript loss (the SFT-time analogue of ECHO-style unmasking — user/tool tokens count toward loss, not just assistant tokens), on the 100-task SWE-bench Verified slice from marin#4898.
Scale-up of SWE-ZERO-10K-Qwen3-1.7B-Base-ECHO-eval and SWE-ZERO-100K-Qwen3-1.7B-Base-ECHO-eval with the same loss-mask change. Companion baseline:… See the full description on the dataset page:
https://huggingface.co/datasets/AlienKevin/SWE-ZERO-1M-Qwen3-1.7B-Base-ECHO-eval.