voicebook-lora-v10-dpo
DPO continue-training from voicebook-lora-v8 with proper reference policy.
Fixes 4 mistakes from v9-DPO:
- Dataset: 1399 pairs (vs 189) — FAIL + WARN + synthetic-broken-PASS
- Continue-from-v8: v8 LoRA loaded as TRAINABLE adapter (not merged-and-wiped),
reference logprobs precomputed with v8 ENABLED
- Schedule: 3 epochs, lr=2e-5, cosine warmup_ratio=0.1
- eval_strategy=no (custom trainer eval was buggy; train metrics monitored)
Final metrics:
- train_loss: 0.0607
- rewards/accuracies (final): 1.0
- rewards/margins (final): ~14-16