An action-conditioned world model for Counter-Strike 2 — feed it keyboard + mouse actions and it
imagines the next frames. Built by adapting MIRA to stream
RekaAI/CS2-10k (63 TB) with no local dataset copy.
Driving the world model with "hold W + sweep the mouse" — every frame above is generated, not real gameplay.
Status: research prototype. It works end-to-end (coherent, action-responsive rollouts) but is
small, blurry, and trained on very little data. Limitations are documented honestly below.
Real vs. generated
Real gameplay (top) vs. world model (bottom)
real vs generated
Codec reconstruction — ground truth ‖ reconstruction
codec reconstruction
Architecture
Stage 1 — Video codec (RAEv2)
Frozen DINOv3-L/16 (facebook/dinov3-vitl16-pretrain-lvd1689m) backbone, layer aggregation
over blocks [11,13,15,17,19,21,23].
Strided-conv bottleneck → 32-channel latent at /32 spatial, /2 temporal.
ViT decoder (width 512, depth 6, patch 16). ~331 M params, mostly frozen.
Stage 2 — Latent world model
Causal flow-matching diffusion transformer over codec latents — 12.5 M trainable params
(hidden 384, 4 layers, 6 heads).
10 keys — W A S D J C R V [ ] (move, jump, crouch, walk, freefall, fire, scope) as a multi-hot
vector, plus continuous mouse deltas.
Upstream MIRA conditions on key presses only (Rocket League is keyboard-only). CS2 is
mouse-driven, so mouse deltas are populated here and consumed by MIRA's existing ActionEncoder
(symlog-normalized + temporally pooled).
Files
File
What
codec_069000.pt
Codec weights (decoder + bottleneck; frozen DINOv3 rebuilt from HF)
world_model_030000.pt
World model weights (12.5 M)
scripts/, src/, configs/
Training / eval / streaming code
patches/dino_hf_backbone.patch
Load gated DINOv3 via transformers instead of Meta's .pth
Training
Codec
World model
Steps
69,000
30,000
Hardware
RTX 5060 Ti (16 GB)
RTX 5060 Ti
Loss
L1 reconstruction
flow-matching diffusion
Final
ema 0.1392
ema 2.41
Trained on the untarred sample/ split: 3 matches / 374 clips / 44 (match, round) groups.
Results
Codec reconstruction: 24.75 dB PSNR (vs 19.7 dB with a random backbone). Recovers room
geometry and recognizable player figures; no fine texture.
Rollouts stay coherent — players persist across the clip, room geometry is stable.
Controllability — mean abs pixel difference of a generated rollout vs an idle-action rollout:
forward vs idle: 0.025 – 0.037
turn vs idle: 0.040
Non-zero and growing over the clip → the model integrates actions over time rather than
ignoring them. This was the go/no-go test, and it passed.
Limitations
Blurry. The /32, 32-channel latent + a depth-6 decoder caps detail (no faces, weapon detail,
floor texture). The paper uses a ViT-XL decoder.
Long-horizon drift. Coherent for several frames past context, then degrades — dissolving to
generic wall texture or darkening. 16-frame clips vs the paper's 80.
Tiny dataset — 3 matches. The dominant constraint; there's little to generalize from.
Tiny model — 12.5 M vs the paper's 1 B. Long-horizon coherence is what capacity buys.
Single-perspective only. MIRA's defining contribution is multiplayer (joint, time-aligned
multi-POV with per-player action attribution). Only the single-player stage is trained here. The
multi-POV data path exists (cs2_stream groups POVs by (match_id, round_number), with ranged
tar-member fetch for the full split) but the joint stage was never trained.
Action semantics unverified. We measure that actions change the output, not that "forward"
moves precisely forward — fidelity is too low to confirm direction visually.
Intended use / out of scope
Research on world models and action-conditioned video prediction. Not a game, an anti-cheat
signal, a player model, or anything deployable. Do not present generated frames as real gameplay.
Attribution & licensing
MIRA — architecture/training code (upstream; this repo ships patches + new modules).