Views
No views yet
| Component | Details |
|---|---|
| Input | Tokenized board [64] |
| Piece vocabulary | 13 IDs: empty, mover pieces, opponent pieces |
| Piece embedding | 13 x 384 |
| Square embedding | Learned absolute 64 x 384 embeddings |
| Board tokens | Piece embedding + square embedding |
| Summary token | Learned 384-dimensional token prepended |
| Transformer | 6 bidirectional transformer layers |
| Attention | 6 heads, 64 dimensions per head |
| Feed-forward | SwiGLU, hidden size 1,040 |
| Output | One 384-dimensional board embedding |
| Causal masking | None |
| Component | Details |
|---|---|
| Current board | Encoded once by the Board Encoder |
| Successor boards | Encoded once per legal move |
| Move embeddings | successor_embedding - current_embedding |
| Candidate count | Typically around 20 legal moves |
| Candidate sequence | Current board copy + all move embeddings |
| Type embeddings | Separate board-token and move-token embeddings |
| Candidate positions | No candidate position embeddings |
| Transformer | 12 bidirectional transformer layers |
| Attention | 6 heads, 64 dimensions per head |
| Feed-forward | SwiGLU, hidden size 1,040 |
| Query source | Original, unaltered current-board embedding |
| Query projection | Bias-free 384 x 384 linear layer |
| Key source | Contextualized move embeddings |
| Key projection | Bias-free 384 x 384 linear layer |
| Normalization | Parameter-free RMSNorm on queries and keys |
| Scoring | Scaled query-key dot product |
| Output | One logit per legal move |
| Padding | Invalid candidates masked to -inf |
python-chess.| Phase | Data and process | Objective and details | Approx. Elo |
|---|---|---|---|
| 1. Supervised policy | Human tournament positions; the recorded human move is the target | Cross-entropy over all legal moves. AdamW; learning rate 4e-4 -> 3e-5; 2,000-step warmup; batch size 32; gradient accumulation 4; BF16 autocast; 0.0 dropout | 1300-1400 |
| 2. Stockfish distillation | The same positions annotated with Stockfish scores for every legal move | KL distillation against the full Stockfish move distribution. Stockfish depth 12; teacher temperature 30; student temperature 1.0; fine-tuned from Phase 1 | 1700-1800 |
| 3. On-policy training | The model generates fresh games while Stockfish plays the opponent; completed rollouts are used once | For the move selected by the model, minimize (model_probability - Stockfish_probability) ** 2. Unchosen moves receive no direct loss. Muon for matrix weights plus AdamW for other parameters; Opponent depth 8; Teacher depth 10; 10,000-step schedule | ~2000 |
checkpoints/policy/on-policy/step-0010000.pt.