This browser-native tournament candidate retrains the 0002 board-snapshot architecture under an exactly matched compute budget. It uses only positions where the eventual winner is to move, from decisive games whose winner has ELO at least 1600. The loser's rating is unrestricted.
Load the public repository in the ChessGPT arena as:
The tournament runner reads browser/manifest.json. The package reconstructs the board from SAN history, scores 4,272 move identities, masks to the runner-supplied legal SAN moves, and returns deterministic argmax SAN.
Frozen January 2026 Lichess standard-rated games only for optimization.
210,285 games scanned, 120,000 accepted, 90,285 filtered, zero invalid.
4,149,869 available winner-side positions; exactly 3,356,140 processed.
Seed 20260729, AdamW, batch size 128, float32 Apple M4 MPS.
11,029,491,768,153,600 ratified training FLOPs, exactly matching model 0002.
No pretrained weights, engine labels, synthetic data, or parent checkpoint.
On 174,404 separately held-out filtered April positions, validation loss was 2.79660, raw top-1 accuracy was 28.486%, legal-masked top-1 accuracy was 30.405%, and the legal-move rate was 100%.
Playing-strength evidence
In a 100-game paired validation match from 50 frozen unfiltered April openings with colors reversed, model 0004 beat model 0002 by 22 wins to 3 with 75 draws, scoring 59.5–40.5. The preregistered prediction was an 80% all-game win rate; the observed win rate was 22%, so the direction was supported but the magnitude was not.
This limited local match is not a precise Elo estimate and does not use the unrevealed official tournament openings. The predicted failure mode—poor play against poor opponents—was not tested.
Integrity
The canonical browser package is 42,585,891 bytes, below the 100,000,000-byte limit. The checkpoint SHA-256 is 98f8ba18a00d35db72b6c58af121614b2c3ff871682475c71b32ed346130d50d. Revision d29db50441c36a109f714b9aafd231fa8e37008c was downloaded cleanly, compared byte-for-byte with the local export, loaded with ONNX Runtime Web 1.27.0, and completed 40 legal SAN self-play plies.
The exact experiment specification, metrics, full loss log, and paired-match artifact are included under training/ and evaluation/.