BDH vs Linear-Attention Baselines — Replication Scaling Study
Independent replication of the scaling experiments from "Burst Denoising Hebbian Neural Networks" (Pathway, arXiv:2509.26507) and a head-to-head comparison of BDH-GPU against GPT-XL, GLA, DeltaNet, Mamba-2 at matched params, same tokens, two GPU replicas (RTX 4080 SUPER, A100-80GB).
Bottom line: BDH reaches consistently lower validation loss at 25M/50M/100M (~0.5–0.8 nats below the best baseline, stable across both GPUs).