Views
No views yet
Honest result: v3 ≈ v2. For deployment, use v2. On the held-out real-defect set the 14 synthetic pairs produced no measurable improvement (the v2→v3 delta is one extra catch and one extra false alarm — noise at n=32), and v3 gave back v2's perfect specificity. The lesson: small synthetic doses don't close the gap; real, objective positives do. v4 acts on that.
| Metric | Base | v2 | v3 |
|---|---|---|---|
| Verdict accuracy | ~72% | 78.1% | 78.1% |
| Positive recall | 87.5% (14/16) | 56.2% (9/16) | 62.5% (10/16) |
| Negative specificity | ~56% | 100% | 93.8% |
| Category match | 56.2% | 43.8% | 43.8% |
| Invalid JSON | 0/32 | 0/32 | 0/32 |
gen_bypass_pairs.py, Drupal-expert-verified) = 526 rows.
QLoRA r=16 on q/k/v/o, batch4+grad-ckpt, MAX_LEN=2048, 3 epochs, lr 2e-4. Full 3-way report with per-item detail
ships in the project repo under docs/eval/dcr-qlora-v3-report.md.