Views
No views yet
Qwen/Qwen2.5-Coder-3B-Instruct. This v3 run uses ReaL's program-analysis detector reward
with DAPO-style token loss and dynamic sampling. The reward is 0.5 * capability_test_fraction + 0.5 * max(0, 1 - 0.3 * detected_vulnerabilities).
Training used seed 42 and the official 655-example training split.
Evaluation used greedy decoding on all 164 official test examples.| Metric | Value |
|---|---|
| Mean reward | 0.500247 |
| Output format pass | 96.95% |
| Syntax pass | 96.95% |
| Capability pass | 24.39% |
| Safety pass | 58.54% |
| Detector clean | 55.49% |
| Detector score | 0.775610 |
| Joint pass | 18.29% |
xw1234gan/seccodeplt-qwen2.5-coder-3b-diff-sft-v2 using alpha=0.5:mixed_logits = 0.5 * pi_theta_logits + 0.5 * anchor_logits