A 57,730-parameter policy that plays a seat in Raifu Wars, a turn-based
strategy game, through the Warrior protocol.
Same architecture and same features as
RaifuWars-RL-ActionScorer. The only
difference is the opponent it trained against: itself, rather than a scripted greedy bot. That
one change is what made it generalise across boards.
Why this model exists
Its predecessor won 55.1% overall, and that number was an average across four maps that hid an
enormous spread:
map
cover tiles
predecessor
this model
Crossroads
0 (0.0%)
68%
69%
Dustbowl
—
80%
not evaluated
Glacier
—
59%
not evaluated
Arboretum
172 (37.2%)
13%
50%
Islands
321 (33.4%)
not evaluated
31%
Crossroads is the only board in the pool with no cover at all. Every policy previously measured
on this game — this net's predecessor and three frontier LLMs — wins there by standing still and
shooting, because four knockouts is an instant win via the KO tier track and nothing blocks line of
sight. Move to a board made of cover and that strategy evaporates:
Chance for one seat of four is 25%. The three LLMs went 0 for 48 on Islands.
On Islands this model averages 0.1 kills per match and still wins 31% — it is not shooting
anything, it is contesting territory. That is a different strategy from the kill-rush every other
policy defaults to, and it is the behaviour the KO-track exploit was masking.
Head-to-head
800 matches in the Hemlock sim, held-out seed, full seat rotation, all four seats drawn from the
policies under test:
policy
win rate
mean return
mean tier
this model
67.4%
15.51
3.65
predecessor (ppo-sim)
25.1%
11.15
3.20
greedy heuristic
5.2%
—
2.69
behaviour cloning
2.2%
7.72
2.69
It gave up nothing in the process: in the single-seat configuration (one learner against three
greedy bots) it scores 80.0% against the predecessor's 75.0% on the same seed.
Training
Initialised from the predecessor, then 7.5 hours of self-play in the
Hemlock reimplementation of the game.
Seven boards, all four seats: Arboretum, Crossroads, Dustbowl, Glacier, Cornfield, Trench Warfare, Twin Rivers.
PPO, lr 5e-5, 8 envs × 256 steps, batch 256.
Use last.pt, not best.pt. Checkpoint selection is by mean return, and under self-play return
is pinned: four copies of one policy always produce exactly one winner, so the terminal term is
fixed at 10/4 = 2.5 and what remains is noise. The best.pt of this run was written 13 minutes in,
never beaten, and loses to last.pt26% to 56% head-to-head. This repo ships last.pt.
Architecture
Two towers and an interaction term. The state is embedded once, each candidate action is embedded,
and the score is their product — so nothing in the network knows how many actions there are, which
is the requirement: the legal set runs from 2 to ~670 between decisions.
state 33 -> 128 -> 64
action 26 -> 128 -> 64
head 192 -> 128 -> 1 softmax over exactly the N offered
value 64 -> 128 -> 1
Scoring rather than classifying makes an illegal action unrepresentable rather than merely
penalised, and lets the same weights run on a 17×21 board and a 27×27 one.
Usage
bash
1python serve.py raifuwars-actionscorer-selfplay.pt --port 89012# then point the game at http://127.0.0.1:8901 via the Warrior protocol
Limitations, stated plainly
Islands is not established as above chance. 5/16 gives a 95% interval of 14–56%, which
includes 25% (p=0.37). What is established is that it beats the LLMs' 0/48 there (p=0.0006) and
the predecessor on Arboretum (8/16 vs 5/39, p=0.006).
Per-map real-game evaluation is 16 matches per board — enough to separate 50% from 13%, not
enough to rank two policies a few points apart.
No terrain input. The 33 state and 26 action features contain nothing describing cover, water
or line of sight; the policy infers position value indirectly. It reaches 50% on a 37%-cover board
regardless, which suggests opponent strength mattered more than perception here — but a terrain
feature set is untested at time of writing.
Pure self-play against the current policy only, with no opponent pool. The arena showed no
forgetting at 7.5 hours; that remains the risk over longer runs.
Only three boards were evaluated in the real game. Dustbowl and Glacier numbers above are the
predecessor's and should not be read as this model's.