Evaluation of a Qwen3-1.7B-Base SFT checkpoint trained on 100K SWE-ZERO trajectories with full-transcript loss (the SFT-time analogue of ECHO-style unmasking — user/tool tokens count toward loss, not just assistant tokens), on the 100-task SWE-bench Verified slice from marin#4898.
Scale-up of AlienKevin/SWE-ZERO-10K-Qwen3-1.7B-Base-ECHO-eval (the same loss-mask change at 10K). Companion to the assistant-only-masked… See the full description on the dataset page:
https://huggingface.co/datasets/AlienKevin/SWE-ZERO-100K-Qwen3-1.7B-Base-ECHO-eval.