Views
No views yet
chess-coach-v6dpo2-4bit-maia, tuned:true). It is a stronger, tier-targeted successor to chess-coach-32b-v6-dpo, built on the chess-coach-32b-v4-qlora SFT base, which remains the full-reproducible-eval reference and the fallback endpoint.chess-coach-v6dpo2-4bit-maia, tuned:true), the best-DPO refinement of the v4 SFT base.unsloth/Qwen3-32B-unsloth-bnb-4bit (the same 4-bit base as v4).| Model (grounded) | tier-policy match | beginner | intermediate | advanced | move-sound | distinct |
|---|---|---|---|---|---|---|
| v4 (SFT-base baseline) | 0.861 | 0.858 | 0.750 | 0.975 | 0.983 | 0.987 |
| v6-dpo | 0.881 | 0.858 | 0.808 | 0.975 | 0.983 | 0.987 |
| v6-dpo2 (this, best DPO) | 0.892 | 0.858 | 0.842 | 0.975 | 0.983 | 0.987 |
select_tier_move rule, a
learnability metric, not certified best teaching.| Resource | Link |
|---|---|
| Shipped model (v4) | chess-coach-32b-v4-qlora |
| Earlier DPO adapter | chess-coach-32b-v6-dpo |
| Preference dataset | chess-coach-v6 |
| Demo | chess-coach-studio |
| Code / GitHub repo | Alpha-AI-Engineering-Khoi/chess-instructor-llm |