Views
No views yet
Qwen/Qwen2.5-Coder-7B-Instruct. This v2 run corrects causal-label alignment and uses the
official ReaL safety-unit-test reward with DAPO-style token loss and dynamic
sampling. Training used seed 42 and the official 655-example training split.
Evaluation used greedy decoding on all 164 official test examples.| Metric | Value |
|---|---|
| Mean reward | 0.511491 |
| Output format pass | 99.39% |
| Syntax pass | 98.17% |
| Capability pass | 39.02% |
| Safety pass | 64.02% |
| Joint pass | 31.71% |
xw1234gan/seccodeplt-qwen2.5-coder-7b-diff-sft-v2 using alpha=0.5:mixed_logits = 0.5 * pi_theta_logits + 0.5 * anchor_logits