Views
No views yet
YWZBrandon/summary-sft-qwen3-4b. Summary generation is
trained with GRPO using DeepSWE next-action preservation and bounded context
length advantages. The KEEP/SUM selector is trained with regret-weighted soft
targets and action-balanced replay. Above the 8,192-token budget, training uses
eight summaries; at or below budget it uses one virtual KEEP reference and
seven summaries.KEEP or SUM verbalizers. For a gate measurement
aligned with training, compare the next-token logits of KEEP and SUM or use
constrained decoding. The summary-generation checkpoint itself is unaffected
by this evaluation-interface issue.