Views
No views yet
No fMRI claim may be read off this model, and this time we know why
The measured BOLD path in this run does integrate the Balloon-Windkessel ODE -- that is what run 4 fixed (ISSUE-008). The haemodynamic likelihood is real here in a way it was not in run 3.It diverges during training, and the full run settled how far.real_bold_nllran 1.99 to 36,472 over 14,600 steps -- a factor of about 18,000 -- whileeeg_nllIMPROVED (1.74 to 1.50), the total loss stayed flat near 1.0, andbold_log_scaleheld at 5.3-5.9. So this is not a variance explosion and the fMRI term is not dominating the mixture: it is the one term getting worse while everything around it gets better. It never plateaus: T4 alone spans 1,530 to 650,815, four orders of magnitude inside one stage. And T5's measured return does not repair it -- that stage grantsbold.*again and ends at 36,472.bold_parcels_coveredheld at full value throughout, so the ODE ran on every parcel for 46 hours and the likelihood it produced is worthless.Four diagnostic arms located the cause: the shared trunk moves out from under the BOLD head.ds002336_realis 5.39% of the source mixture and is outvoted 17.6 : 1 by the EEG-like sources, so the latent state converges on what they want. Freezing the five Balloon parameters changes nothing; freezing the trunk makes the BOLD likelihood improve.ds002336_realappears incontributed_sourcesand that is accurate: its BOLD channel contributed a gradient. Whether that gradient carried information is a separate question, and this run's leave-one-source-out answers it: removing the corpus made measured EEG prediction worse (+0.0010), so the information went into the shared state and does not come back out through the BOLD head. Both things are true at once and neither licenses an fMRI claim.No fMRI, haemodynamic or neurovascular claim about this artifact is supported. ISSUE-016 is open; the remedy is unshared capacity for the slow modality, which is a later run's design, not a caveat on this one.No inference or parameter-recovery claim may be read off this model. ISSUE-012 is open and this run MEASURED it rather than inheriting it. The learning-rate repair worked -- bestlog_GR^2 0.284, against ~0 on every parameter in run 3, so the flow now reads its conditioning -- and it overshot: worstposterior_z_sd59.3 where a calibrated posterior sits near 1.0, SBC KS p_min 1.03e-147, coverage MAE 0.203. The posterior narrowed far more than its accuracy earned, so it is confidently wrong rather than uninformative. 4 of 6 parameters still explain no variance (log_velocity,ei_global,ei_gradient,drive). Run 3's posterior was uninformative and honest; this one is partly informative and overconfident. Neither supports inference.No individualisation or personalisation claim may be read off this model, and this run MEASURED that rather than inheriting it.session_individualisationscored 75 participants over 1500 held-out second-night windows -- the same people on both sides of a SESSION split, which is the only arrangement on which a person effect is measurable at all. Held-out session NLL 2.0436 [2.0048, 2.0897], bootstrapped over participants rather than windows.That score is not the finding. The between-participant spread of the applied theta shift is 0.000706 -- 0.67% of the scale the model allocated for that effect. 30 of the scored person-effect rows are exactly zero. The individualiser applied essentially nothing on a split built specifically to let it apply something, which is the falsifier this evaluation declares for the capability. Earlier runs reported individualisation as unmeasurable on a participant-disjoint split; this one built the split, trained the effect and measured it, and the effect is a fraction of a percent of its own scale.The held-out NLL is reported because withholding a measured number is its own distortion. It answers a different question than it appears to: what separates an individualised model from the population model here is the shift, not the score.Which sources earn their place: leave-one-source-out on the MEASURED holdout
Each arm drops one source family, retrains 200 steps, and is scored on the same held-out participants the headline rests on. Positive delta means removing the family made measured prediction WORSE, i.e. the family contributed.2 of 10 families contributed on the measured holdout; 8 showed negative transfer (anatomical_prior,ds000117_behaviour,ds000117_real,ds004024_perturb,ds004024_rest_real,montage_calibration,sim_wholebrain,sleepedf_real). Largest positive deltas:eegmmidb_real+0.0144,ds002336_real+0.0010.The simulated-holdout arm of the same run is retained for comparability with earlier runs and is NOT the result: 5 of 10 families come back as negative transfer there. Scoring a measured source against the simulator asks whether dropping it helps the model fit the simulator, which is a different question from the one above and is not evidence about the source.The EEG lead field remains an analytic sphere, not a head model, so no source-localisation claim is available.
Read this first: this checkpoint loses to copying the last observed sample forward. It is published as a negative result and as a reference artifact for others, not as a working model. If you are looking for a brain-dynamics model that works, this is not it.
This model never saw measured data during training, and four other training mechanisms were silently off. This run renamed its training stages; six gates in the trainer match on the previous run's stage names, and five of the six therefore gave the wrong answer:
gate decides result for this run admit measured sources refused -- no gradient was ever taken on real EEG per-stage gradient allowlist wildcard -- no restriction applied boundary randomisation of sim inputs off haemodynamic state in the rollout off build the individualizer off -- the stage named for it ran ordinary training admit simulated sources admitted, and correct only by accident Nothing raised, because every gate fails toward permissive: an unmatched stage name means "no restriction" rather than "unknown stage". The scores below are therefore a simulation-to-measurement transfer result from a partially configured trainer -- not a held-out-performance result -- and a stage named for measurement is not evidence that measurement occurred. Seereports/RUN2.mdsection 2b.
The comparison below flatters this model, and the correction is not applied to the numbers.evaluate.pyscores SC-WBD ony = target / s, wheresis each window's own standard deviation, with the Jacobian folded into the log-variance. Every baseline is scored on the raw target. The algebra is exact and model-independent:NLL_scaled = NLL_raw - log s, so the two sides are different random variables and SC-WBD's figure is the smaller one.Measured on this test fold,mean(log s) = 0.5694nats -- against a spread of 0.035 nats across the three non-trivial baselines, so the offset is roughly 17x the entire spread it is being compared against. In the baselines' units SC-WBD's NLL is approximately 3.75 rather than the 3.179 tabulated, and the gap to the best baseline is about 1.70 nats rather than 1.13. MSE is off by1/s^2and does not cancel.The rescale is harmless during training --sdoes not depend on the parameters -- and is a pure unearned advantage at evaluation. No verdict changes: every paired interval already excluded zero and the correction moves all of them further from SC-WBD. The numbers are left as measured rather than silently adjusted, because re-scoring both sides on the raw target is the fix and arithmetic on published figures is not.
No baseline beats scwbd-004, and scwbd-004 is not shown to beat ar16, var4 (3 of 5 comparators separated) on the paired participant-clustered 95% interval of the per-window NLL difference
| arm | NLL | 95% CI | MSE | params |
|---|---|---|---|---|
scwbd-004 ← | 2.0244 | [1.9750, 2.0903] | 3.9822 | 26,304,657 |
ar16 | 2.0345 | [1.9696, 2.1349] | 3.9856 | 5,184 |
var4 | 2.0371 | [1.9689, 2.1439] | 3.9622 | 20,544 |
population_gaussian | 2.0589 | [1.9975, 2.1526] | 4.1505 | 2,208 |
persistence | 2.3274 | [2.2546, 2.4305] | 7.3341 | 4,096 |
dense_neural | 3.4062 | [3.2229, 3.6615] | 14.3610 | 26,296,869 |
| vs | Δ NLL | 95% CI | excludes zero |
|---|---|---|---|
ar16 | -0.0100 | [-0.0480, 0.0144] | False |
var4 | -0.0127 | [-0.0561, 0.0146] | False |
population_gaussian | -0.0345 | [-0.0654, -0.0133] | True |
persistence | -0.3030 | [-0.3427, -0.2691] | True |
dense_neural | -1.3818 | [-1.5750, -1.2380] | True |
eeg.log_noise) sets the predictive variance, was left to SGD instead of its closed-form optimum, and ended up uniformly overconfident. The scalar cannot represent horizon-dependence at all; the baselines' variance can.Schaefer400x7 in fsLR/32k, subcortex Aseg14T, 10 declared sources (enigma_hcp_sc, hansen_receptors, hcps1200_maps, hill2010, margulies2016, neuromaps, raichle_metabolism, schaefer2018, sydnor2021, tian2020); is_biological = True. Full provenance, including every licence and citation, is carried inside the checkpoint under extra.anatomy and in reports/anatomy_prior.md.analytic_sphere_fallback, individual head model: False.gradient_permission. They carry no parameters in this checkpoint's parameter report, so the share of the model that could not receive a gradient is 0.0% -- this is a completeness note about the cards, not a finding about the weights. Computed from the source cards at b20f368, the commit this checkpoint records. That commit is recorded with a -dirty suffix, so the tree that trained carried uncommitted changes and the cards it used may differ from the cards at the commit.subject_specific_ar is not in the table. DROPPED from this table under baseline protocol v2 (ISSUE-013). Refusal R10 makes the fit and score participant sets disjoint, so 100% of scored windows routed to the pooled fallback and the row was bit-for-bit ar16 -- a duplicate carrying the name of the hardest baseline the thesis names. Protocol v1 (runs 1-3) reported it; those numbers stand as ar16's. The quantity it was supposed to measure is measured instead by within_participant_holdout, on a within-participant temporal split, and is NOT comparable with the rows here: it has seen the scored participant and every row here has not.non-commercial: yes; share-alike: yes; attribution: required; redistribution: unknown; SHARE-ALIKE IN FORCE: derivative works must be released under the same licence; 1 source(s) with UNKNOWN licence (montage_calibration) — unknown is not permissivescwbd.release.licence.union_of, not asserted here.ATTRIBUTION
checkpoint: scwbd-004
============================================================
DATASET INPUTS (7)
ds000117 (dataset)
cite: Wakeman DG, Henson RN (2015). A multi-subject, multi-modal human neuroimaging dataset. Scientific Data 2:150001, doi:10.1038/sdata.2015.1. OpenNeuro dataset ds000117 v1.1.0, doi:10.18112/openneuro.ds000117.v1.1.0.
licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
doi: 10.18112/openneuro.ds000117.v1.1.0
from: scwbd/sources/cards/ds000117.yaml
ds000117 (dataset)
cite: Wakeman DG, Henson RN (2015). A multi-subject, multi-modal human neuroimaging dataset. Scientific Data 2:150001, doi:10.1038/sdata.2015.1. OpenNeuro dataset ds000117 v1.1.0, doi:10.18112/openneuro.ds000117.v1.1.0.
licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
doi: 10.18112/openneuro.ds000117.v1.1.0
from: scwbd/sources/cards/ds000117.yaml
ds002336 (dataset)
cite: Lioi G, Cury C, Perronnet L, Mano M, Bannier E, Lecuyer A, Barillot C (2020). Simultaneous MRI-EEG during a motor imagery neurofeedback task: an open access brain imaging dataset for multi-modal data integration. Scientific Data 7:173, doi:10.1038/s41597-020-0498-3. OpenNeuro dataset ds002336 v2.0.2, doi:10.18112/openneuro.ds002336.v2.0.2. Paradigm: Perronnet L et al. (2017), Front Hum Neurosci 11:193.
licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
doi: 10.18112/openneuro.ds002336.v2.0.2
from: scwbd/sources/cards/ds002336.yaml
ds004024 (dataset)
cite: Hernandez Pavon JC, Schneider Garces N, Begnoche JP, Miller LE, Raij T (2022). OpenNeuro dataset ds004024, doi:10.18112/openneuro.ds004024.v1.0.0. Cortico-cortical paired associative stimulation (ccPAS) with bi-focal MRI-navigated TMS-EEG of left and right M1.
licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
doi: 10.18112/openneuro.ds004024.v1.0.0
from: scwbd/sources/cards/ds004024.yaml
ds004024 (dataset)
cite: Hernandez Pavon JC, Schneider Garces N, Begnoche JP, Miller LE, Raij T (2022). OpenNeuro dataset ds004024, doi:10.18112/openneuro.ds004024.v1.0.0. Cortico-cortical paired associative stimulation (ccPAS) with bi-focal MRI-navigated TMS-EEG of left and right M1.
licence: [CC0-1.0] CC0 1.0 Universal (public domain dedication)
doi: 10.18112/openneuro.ds004024.v1.0.0
from: scwbd/sources/cards/ds004024.yaml
eegmmidb (dataset)
cite: Schalk G, McFarland DJ, Hinterberger T, Birbaumer N, Wolpaw JR (2004). BCI2000: A General-Purpose Brain-Computer Interface (BCI) System. IEEE Trans Biomed Eng 51(6):1034-1043. Dataset: Schalk G (2009), EEG Motor Movement/Imagery Dataset (version 1.0.0), PhysioNet, RRID:SCR_007345, https://doi.org/10.13026/C28G6P
licence: [ODC-By-1.0] Open Data Commons Attribution License v1.0 (ODC-By 1.0)
doi: 10.13026/C28G6P
from: scwbd/sources/cards/eegmmidb.yaml
sleep-edfx (dataset)
cite: Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL (2000). Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG. IEEE Trans Biomed Eng 47(9):1185-1194. Dataset: Kemp B, Zwinderman AH, Tuk B, Kamphuisen HAC, Oberye JJL. Sleep-EDF Database Expanded (version 1.0.0), PhysioNet, https://doi.org/10.13026/C2X676
licence: [ODC-By-1.0] Open Data Commons Attribution License v1.0 (ODC-By 1.0)
doi: 10.13026/C2X676
from: scwbd/sources/cards/sleep-edfx.yamlscwbd/release/publish.py. None of them is typed into the card generator. The sources:reports/training/evaluation_run4.json — every score, CI, parameter count and split sizeconfigs/scwbd_001_beta.yaml — the training mixturescwbd/sources/cards/*.yaml — dataset citations and licencesreports/scope_gap.md, reports/training/p0_variance_channel.md — the two diagnoses, stated in prose above