Views
No views yet
config.yaml, dataset_statistics.json, the selected checkpoint(s), and an EVAL.md with the
exact run command. For experiments with more than one saved checkpoint we keep the final step
and one intermediate step; single-checkpoint experiments keep the only one.EVAL.md.aim_cos = eval mouse-aim cosine similarity (the aim quality metric; ~0 means aim did not
learn). key_f1 = eval h0 key macro-F1. attack_f1 = eval firing F1.| folder | checkpoint | step | aim_cos | key_f1 | attack_f1 |
|---|---|---|---|---|---|
csbc | steps_5000.pt (best eval) | 5000 | 0.00 | 0.285 | 0.156 |
csbc | steps_15000.pt (final) | 15000 | 0.00 | 0.265 | 0.179 |
nitrogen_starvla | steps_30000.pt (mid) | 30000 | 0.00 | 0.179 | 0.058 |
nitrogen_starvla | final_model.pt (final, complete 60k) | 60000 | 0.01 | 0.191 | 0.065 |
aimflow | steps_5000.pt (only) | 5000 | 0.475 ⚠️ | 0.143 | 0.042 |
aimflow_v2 | steps_10000.pt (mid) | 10000 | -0.01 | 0.151 | 0.027 |
aimflow_v2 | steps_15000.pt (final) | 15000 | 0.01 | 0.158 | 0.021 |
aimflow_v3 | steps_30000.pt (mid) | 30000 | -0.03 | 0.182 | 0.117 |
aimflow_v3 | steps_55000.pt (final) | 55000 | 0.04 | 0.157 | 0.025 |
aimflow (v1) aim_cos ≈ 0.475 is a metric artifact, not real aim: v1 used a different mouse
normalization, so its aim_cos is NOT comparable to v2/v3 (which read ~0). Across all experiments here,
aim is not convincingly learned — treat it as an open question to verify in the real environment.csbc peaks early (step 5000) then overfits on keys; steps_15000 is more trained but slightly worse.nitrogen_starvla/final_model.pt is the only fully-completed (60k-step) run.