Views
No views yet
yoheikobashi/ptcg-qwen3-4b-cardfirst-v40, trained with DPO on
playout-gated preference pairs mined from the policy's own low-margin decisions
(24 playouts/label, +-4pt ship/revert gates). dpo_r8/ is the final full-roster round;
lora_<deck>_r*/ are per-deck rounds.