Daily spatio-temporal datacube for wildfire ignition and cause prediction over Washington State. Ten geospatial products spanning terrain, fuels, weather, human activity, lightning, and fire history are aggregated onto a single daily 2km $\times$ 2km grid spanning 2003-2020 continuously.
Full daily coverage from 2003-2020, supervised over the May-October fire season (months that carry active fire activity).
38 input channels (25 from source processors, 13 derived), normalized, resampled and interpolated into trainable feature layers.
Custom derived features: Per-cause ignition KDE's, 3x3 cell 7d rolling fire occurrence, NDVI anomalies, 2 and 5-day cumulative precipitation, 100 and 1000-hr dead fuel moisture, decayed lightning load, Fosberg FWI.
Circular quantities (N/S and E/W aspect and wind-direction components, day-of-year) decomposed into orthogonal components, so no channel carries 0/360 discontinuities.
Labels for fire ignition time, cause, KDE by cause, 3x3 rolling fire occurrence, and months since last burn.
Mask layers for water and active fire, for ignition prediction.
Built by Tanner O'Rourke as part of FireFusion, independent research conducted on multi-source wildfire ignition modeling. The full pipeline, from raw-source extraction to model training, is at FireFusionNet. If this work is useful to you, a star on the repository helps it reach more people. Feel free to reach out!
bibtex
1@misc{orourke2026firefusion,
2 author = {O'Rourke, Tanner},
3 title = {FireFusion: multi-source datacubes and a spatiotemporal ConvFormer for wildfire ignition prediction},
4 year = {2026},
5 url = {https://github.com/tannerorourke/FireFusionNet}
6}
Files
Artifact
Definition
dataset.zarr
Primary artifact. Only deterministic functions of the raw sources. every supervised day, split-agnostic.
dataset_manifest.json
grid, channel list, and every deterministic transform
train.zarr / eval.zarr / test.zarr
the suggested splits, described below
manifest.json
split years, channel order, and the train-fitted statistics applied to those splits
dataset.zarr
The primary artifact, cut prior to compilation; Every transform with a data-estimated parameter (any z-score, min-max, or scale) is deferred out of it and into the split stage, so it carries no knowledge of where the split boundaries fall. The time axis carries the same off-season gaps as the splits, so the contiguous-block check shown for those applies here unchanged.
Differences from the splits store:
No statistical normalization. Only the deterministic steps (clip, log1p, to_sin, per_area) are applied, so values sit in native units after those steps. Every z-score and min-max is fit at compile time against whichever years you designate as training, which is what keeps this store free of split-boundary leakage.
Four cause classes, not three.ign_next_cause runs 0-3 with DEBRIS held separate from INDUSTRIAL, and dataset_manifest.json records n_cause_classes: 4. See Labels.
46 channel variables versus the splits' 38. The extra eight are raw layers the derived channels are built from, kept for provenance and dropped by compile: lf_aspect and wind_dir before their component decomposition, modis_ndvi before the anomaly, plus modis_water_mask, rh_max, usfs_burn_occ, usfs_burn_cause, and usfs_perimeter.
Features are held as named variables rather than a stacked tensor, so a channel loads independent of the other 45.
Schema
Group
Variables
Dims
dtype
channels
45
(time, y, x)
float32 x42, uint8 x3
usfs_burn_cause
1
(time, burn_cause, y, x)
uint8
labels
2
(time, y, x)
int8
masks
3
(time, y, x)
uint8
Coordinates are time (datetime64[ns], 3312 days), y and x (float64, UTM metres), and burn_cause (the four cause names). Every array is chunked (16, 204, 217), 16 days at full spatial extent. Store size 11 GB.
Loading
python
1ds = xr.open_zarr("dataset.zarr")# 3312 days, no split boundary23vpd = ds["vpd_max"]# (3312, 204, 217) float324window = ds[["temp_max","wind_mph","lightning_load"]].isel(time=slice(0,10))5y, cause = ds["ign_next"], ds["ign_next_cause"]67# stack an arbitrary channel set into a model-shaped tensor8chans =["temp_max","vpd_max","lf_elevation"]9X = ds[chans].to_array("channel").transpose("time","channel","y","x")
To cut your own splits, set train_yrs / eval_yrs / test_yrs in fire_fusion/config/dataset_config.py and run compile against this store. Every statistic refits on the years you name:
Train-ready splits. Each is a zarr store holding X with shape (time, channel, y, x) in float32, stacked tensor along with labels and masks. The task is forecasting, these suggested splits are chronological and balanced positive label counts with training data size. Note that any random sampler will leak future weather in training.
Split
Years
Days
Size
train (train.zarr)
2003-2016
2576
8.9 GB
eval (eval.zarr)
2017-2018
368
1.3 GB
test (test.zarr)
2019-2020
368
1.4 GB
The splits are cut for comparability and not the only choice. Cut your own with the compile stage of the FireFusionNet repo, which refits every statistic on whatever you designate as training data.
Schema
Variable
Dims
dtype
X
(time, channel, y, x)
float32
ign_next, ign_next_cause
(time, y, x)
int8
land_mask, no_act_fire_mask, valid_cause_mask
(time, y, x)
uint8
channel is a labelled coordinate carrying the names in manifest.json, so ds["X"].sel(channel="temp_max") resolves without an index lookup. time is datetime64[ns], and y and x are float64 UTM metres. X is chunked (16, 38, 204, 217), labels and masks (64, 204, 217).
Loading
python
1import numpy as np
2import xarray as xr
34ds = xr.open_zarr("train.zarr")# local path, or an fsspec URL5x = ds["X"].isel(time=slice(0,10))# (10, 38, 204, 217)6y = ds["ign_next"].isel(time=9)78# the time axis skips the off-season, so build windows within a contiguous block9days = np.asarray(ds.indexes["time"], dtype="datetime64[D]")10block = np.concatenate([[0], np.cumsum(np.diff(days).astype(int)!=1)])11# a length-W window at t is valid iff block[t] == block[t + W - 1]
Note on Streaming
If training or compiling the data, download first. Streaming is only recommended for inspection; split X chunks are (16, 38, 204, 217), about 108 MB each decompressed and ~55 MB on the wire. A 10-day window spans one or two of them, so a train epoch of ~2450 windows moves on the order of 200 GB, roughly 24x the cost of downloading the 8.9 GB split once.
Grid Details
Stat
Value
CRS
EPSG:32610 (UTM Zone 10N)
resolution
2000 m
grid (y, x)
204 x 217
latitude
45.5 to 49.0
longitude
-122.5 to -117.0
season
May 1 - Oct 31, each year 2003-2020
supervised days / year
184
supervised days total
3312
Time axis
Only in-season days ship, so the time axis is not contiguous: within a year it runs May 1 to Oct 31 day-by-day, then jumps to the next May. The manifests record time.season_months = [5, 10] and time.contiguous = false.
Halo days
Each year is built from a single block running March 22 to November 10, deliberately wider than the window. This halo -a 40-day lead and a 10-day trail - lets temporal derivations enter the supervised window with real history instead of restarting at zero. Every backward-looking channel (decayed lightning load, 2 and 5-day cumulative precipitation, the per-cause ignition KDEs) is computed on the wider index, and the halo is dropped before any normalization statistic or class balance is taken, so those describe exactly the days that ship. The 40-day lead is sized by the longest backward operator in the pipeline, the lightning-load IIR, which decays below 0.1% there.
Halo days are never supervised and never scored. They survive only in the build-time staging cube (cube.zarr), which is not distributed, and are absent from every artifact listed above. A sliding-window loader must still avoid building any window that straddles the year-to-year gap. See Loading.
State that legitimately spans years decays by elapsed time, not by index position. The per-cause ignition KDEs apply a 365-day half-life to the true day count between consecutive entries, so a multi-year prior crosses the roughly 4.5-month off-season gap correctly attenuated instead of stepping across it as a single day.
234 days are extracted per year against the 184 that are supervised. Over 2003-2020 that is 4212 days extracted and 3312 published.
Labels
Name
dtype
Meaning
ign_next
int8
1: if a clear cell burns within the next 7 days
ign_next_cause
int8
cause id of the earliest such ignition, else -1
ID
Cause
Train positives
0
NATURAL_LIGHTNING
15674
1
HUMAN
6789
2
INDUSTRIAL (includes debris)
813
Note: dataset.zarr carries a fourth class DEBRIS. The splits compileDEBRIS into INDUSTRIAL together, since both are a few hundred cases, versus lightning and human causes which are the overwhelming majority of cases. Regroup them by re-running compile against the published cube; n_cause_classes in each manifest is authoritative for the store beside it.
Masks
Name
Meaning
land_mask
1 on land, derived from the MODIS water flag
no_act_fire_mask
1 where the cell is not already burning
valid_cause_mask
1 where ign_next_cause carries a usable label
Ignition is heavily imbalanced: ign_pos_weight = 2683.80 on the train split, so positives are ~3.7e-4 of supervised cell-days. Restricting to the fire season removes winter cell-days that are near-uniformly negative, so this is a fire-season base rate and not an annual one. Losses should apply land_mask and no_act_fire_mask; the cause head should additionally apply valid_cause_mask.
Channels
Deterministic steps (clip, log1p, to_sin, per_area) are already applied in dataset.zarr and recorded in dataset_manifest.json. Statistical steps
(z_score, minmax) are fit on the train years only and applied in the split stores, recorded in manifest.json. Both are listed together below in the order they compose.
In-depth extraction and feature details can be found in the repo's SOURCING.md.
#
Channel
Normalization
0
canopy_cover_pct
clip
1
d_to_road
clip -> log1p -> z_score
2
dead_fmo_1000hr
z_score
3
dead_fmo_100hr
z_score
4
dewpoint
z_score
5
doy_sin
(none)
6
fire_spatial_roll
log1p -> z_score
7
fosberg_fwi
z_score
8
frac_imp_surface
clip
9
kde_debris
per_area -> z_score
10
kde_human
per_area -> z_score
11
kde_industrial
per_area -> z_score
12
kde_natural_lightning
per_area -> z_score
13
lf_aspect_ew
(none)
14
lf_aspect_ns
(none)
15
lf_elevation
z_score
16
lf_slope
z_score
17
lightning_load
log1p -> z_score
18
lightning_strikes
log1p -> z_score
19
modis_lai
clip -> z_score
20
modis_months_since_last_burn
log1p -> minmax
21
ndvi_anomaly
clip -> z_score
22
pop_density
clip -> log1p -> z_score
23
precip_2d
log1p -> z_score
24
precip_5d
log1p -> z_score
25
precip_mm
log1p -> z_score
26
rel_humidity
z_score
27
temp_avg
z_score
28
temp_max
z_score
29
temp_min
z_score
30
usda_dist_to_wui_km
z_score
31
usda_hs_density_km2
log1p -> z_score
32
usda_wui_index
z_score
33
vpd_max
clip -> log1p -> z_score
34
vpd_min
clip -> log1p -> z_score
35
wind_dir_ew
(none)
36
wind_dir_ns
(none)
37
wind_mph
clip -> log1p -> z_score
The four kde_* channels are fire-history kernel density estimators in events per km squared, with a 20 km smoothing radius and a 365-day decay half-life. The per_area step is what puts them in those units; the raw accumulator is mass per cell and would rescale with cell size.
ndvi_anomaly is NDVI minus its day-of-year climatology, averaged over the train years only so no held-out day contributes to the mean it is measured against.
Sources
dataset_manifest.json and manifest.json are the authoritative record of channel order, normalization steps, grid geometry, and class balance. Read them rather than hardcoding.
MODIS MOD13Q1 / MCD15A2H / MCD64A1: NDVI, leaf area index, burn history, water mask
LANDFIRE: elevation, slope, aspect
NLCD: canopy cover, impervious surface
USFS FOD and fire perimeters: ignition and cause labels
NOAA NCEI SWDI: lightning
GPW: population
USDA WUI: wildland-urban interface
US Census TIGER: roads vectors
Fire Fusion collection
All FireFusion datasets carry identical channels, labels, masks, and split years. Three are statewide, cascades500 covers only the Eastern Cascades sub region.
Note: Class balance is not identical; Ignition prevalence per cell rises with cell size. (e.g., ign_pos_weight is 944.85 at 4 km vs. 2683.80 at 2 km). See manifest for exact values.