Zeno Divergent v0.15.0
1STATUS
2v0.15.0 · PUBLISHED · weights downloadable on `main`
3gate: PROMOTED (criteria v0.15.0) · claim scope: reference lane only · Tier B (operational)
4research artifact — not a production safety system
1ACTIVE RESEARCH CANDIDATE (documentation only — no weights published)
2v0.16-rc1 · run v016-rc1-20260906-r1 · RUNNING on 1x A100-80GB
3forecast-aware, two pinned sources: Flatplanet Kp archive (fp-kp-20260906T191207Z, 80,750 rows)
4and causal Hubverse ledger (hubverse-causal-20260906T204727Z, 1,855 declared model-periods)
5arms: D-self 12,136,711 params vs peer-aware D-rel 13,346,696 params (never summed)
6NO RESULTS YET · NOT RELEASED · v0.15.0 remains the published artifact on `main`
What this model is
A small model whose subject is other models. Zeno Divergent reads the behaviour of named forecasting models — their gaps, errors, baseline margins, regime statistics — and estimates when such a model is approaching the boundary of its validated reliability. It does not forecast the world.
Research direction: a model of models, i.e. a shared surprise state that transfers across unrelated domains and that improves when peer models are visible. The first half is measured here. The second half (relational peer value) is not demonstrated and is the subject of v0.16.
Current claim scope
| Scope | Status |
|---|
| Reference lane (Zeno in-house watched forecasters) | Tier B — operational predictive signal. The only published claim. |
| Seismic | Narrow signal, not claimed; significant_event_ahead calibration fails (test ECE 0.164). |
| Climate, hydrology, markets, space weather | Measured and passed the bar as down-weighted evidence lanes; outside claim scope, shipped as labelled candidates. |
| Hydrology vs. gradient boosting | Beat GBM on 4 of 5 targets — the first measured Tier A representation result in any cycle, on an evidence lane, therefore not a claim. |
| Agriculture, Celo on-chain | Measured and refused. |
Tier A (learned state beats a strong tabular baseline on the same features) is unmet inside claim scope.
Supports / does not support
Supports:
- Estimating, for a watched forecaster inside the reference lane, the probability of a validity breach, an error-regime shift or a baseline crossover in the next window.
- Calibrated abstention on those targets (worst claimed-target test ECE 0.033).
Does not support:
- General model-failure prediction for arbitrary models.
- Relational peer value — that peers improve the call on a given model.
- Transfer to unseen model families.
- Causal attribution of why a model degrades.
- Any production safety guarantee.
Evidence
Reference lane, chronological test split, 1,566 entities / 99,449 rows. AUROC against the operational rule baseline, with logistic and gradient-boosted baselines on identical features.
| Target | AUROC | rules | logistic | GBM | ECE | lift vs rules | claimed |
|---|
error_regime_shift | 0.680 | 0.518 | 0.693 | 0.687 | 0.012 | +0.163 (sig.) | yes |
validity_breach_ahead | 0.679 | 0.549 | 0.682 | 0.713 | 0.033 | +0.130 (sig.) | yes |
baseline_crossover_ahead | 0.567 | 0.471 | 0.659 | 0.626 | 0.018 | +0.096 (sig.) | yes |
state_transition | 0.576 | 0.610 | 0.530 | 0.682 | 0.032 | −0.036 (n.s.) | no |
Read this honestly: the claimed heads beat the deployed rule baseline with confidence intervals excluding zero, and they are well calibrated — but they do not beat gradient boosting on the same features. That is precisely why the claim is Tier B and not Tier A.
Evidence lanes (not claimed): climate 0.78–0.88 AUROC across five targets; hydrology 0.64–0.78 with four GBM wins; markets 0.63–0.68; space weather 0.67–0.71. Refused: agriculture (no target cleared, ECE unmeasurable), Celo (lift +0.052/+0.058, CI includes zero).
Leave-one-domain-out transfer (trunk + latent trained on the other domains, frozen, only held-out adapter/heads fitted) is reported per lane in report.json. It is reported, never required.
Architecture
1per-domain adapter → shared dilated causal trunk (d_model 512, 12 steps, dilations 1/2/4)
2 → shared surprise latent z^S (160)
3 → per-domain heads
- Parameters: 6,623,800 total · 5,363,433 shared trunk + latent; the rest are per-domain adapters and heads.
- Architecture version
surprise-arch/0.6.0; counter-expectation branch, domain-adversarial head (8 domains), baseline-residual head, dropout 0.1, context dim 4.
Training data
Trained from a durable append-only mirror, 2,233,367 rows, plus per-lane feature caches. Zero live partner API calls — the partner platforms behind several original lanes are shut down, so the cycle is reproducible without them.
Real data providers are named; transport APIs are never counted as sources.
| Lane | Real providers | Weight | Rows / entities (test scope) |
|---|
| Watched models (reference) | Zeno in-house reference forecast ledgers + Flatplanet peer receipts (263,139) | 1.00 gated | 99,449 · 1,566 |
| Seismic | USGS FDSN event catalogue | 1.00 gated | 49,123 · 211 |
| Climate | NOAA GHCN-Daily stations | 0.50 evidence | 148,785 · 182 |
| Hydrology | USGS Water Services daily values | 0.50 evidence | 85,120 · 320 |
| Markets | Coinbase Exchange public market data | 0.50 evidence | 52,822 · 357 |
| Space weather | GFZ Potsdam Kp/ap/F10.7 (definitive) | 0.50 evidence | 12,960 · 300 |
| Agriculture | National agro-meteorological services | 0.25 evidence | 620 · 31 |
| Celo on-chain | Celo public JSON-RPC block sampling | 0.35 evidence | 2,500 · 250 |
Training-mix law: claim-isolated. Evidence weights change how much a scorable row counts inside its class; they never move a domain across the claim line (evidence-weight/0.1.0, hash b352a5821f224468).
Evaluation protocol
- Chronological, availability-aware splits keyed on
available_at; horizon selection on calibration only, never on test.
- Baselines on identical features: operational rules, EWMA-z, logistic regression, gradient-boosted trees. Beating persistence or climatology alone never earns a claim.
- Promotion bar (declared before the run, criteria
v0.15.0): ≥1 well-posed target beating the rule baseline by +0.03 AUROC with a CI excluding zero, worst calibrated ECE ≤ 0.15, non-negative selective-accuracy gain, ≥10 scorable entities, all leakage tests passing.
- Tier labels are not comparable across criteria versions.
Refused claims
agriculture and celo were measured and refused. state_transition is measured and not claimed. Seismic ships with its calibration failure disclosed. Relational peer value was measured in the v1 continuous protocol and not supported (validation MAE skill vs persistence −0.042; binary breach heads retired for having fewer than 25 training positives). Refusal is a published outcome, not an omission.
Quick start
1import json, torch
2from huggingface_hub import hf_hub_download
3from safetensors.torch import load_file
4
5repo = "mbarbosa1/zeno-divergent-v1"
6cfg = json.load(open(hf_hub_download(repo, "config.json", revision="v0.15.0")))
7w = load_file(hf_hub_download(repo, "model.safetensors", revision="v0.15.0"))
Reference implementation (model, data, losses, baselines, evaluation) is in reference_implementation/. Pin revision="v0.15.0"; main moves.
Reproduce
| Field | Value |
|---|
| Run id | v015-a100-20260905 |
| Job id | 6a9c0559259f8e97255e2bd1 |
| Hardware | 1× NVIDIA A100 (Hugging Face Jobs), 20 epochs |
| Seed | 1 |
| Data snapshot | durable-mirror-2233367 |
| Architecture version | surprise-arch/0.6.0 |
| Sequence version | surprise-sequence/0.4.0 |
| Evidence-weight version | evidence-weight/0.1.0 (b352a5821f224468) |
| Gate criteria version | v0.15.0 |
model.safetensors sha256 | 8b448fb571bd73756bb84162abec0e5945b80ee6099400784455b1d61f9d17f2 |
report.json sha256 | c8735a40fa170ebafcba0f818b33cfe247e0eb070af35e0be93b6ef30cd7106f |
config.json sha256 | 289b39cc984b821ca8384d3a0dab626d9ddf647e8f2608c6dce58aee401db1c3 |
Where this card and report.json disagree, the report wins.
Browser scorer: the in-browser ONNX export under browser/v0.10.0/ still runs v0.10.0 heads and does not reproduce v0.15.0 numbers.
Release history
Each release note is immutable and carries hypothesis, data frozen, architecture, baselines, result, failures, claims allowed and refused, artifact hashes.
| Release | Note | Claim |
|---|
| v0.15.0 | releases/v0.15.md | Reference Tier B; hydrology beat GBM on an evidence lane |
| v0.14.0 | releases/v0.14.md | Reference + climate Tier B; seismic Tier C |
| v0.11.0 | releases/v0.11.md | Climate Tier B; watched-model lane Tier C; negative cross-model ablation |
| v0.16a | experiments/v0.16a-predeclaration.md | Predeclared, untrained |
Citation
1@software{zeno_divergent_2026,
2 title = {Zeno Divergent v0.15.0: a small model of model validity},
3 year = {2026},
4 url = {https://huggingface.co/mbarbosa1/zeno-divergent-v1},
5 note = {Run v015-a100-20260905, gate criteria v0.15.0}
6}
License CC-BY-4.0. Entities are sensors, sites, instruments, processes and models — never bind an entity to a natural person.