JumpReLU Sparse Autoencoders trained on every residual stream layer of google/gemma-4-E2B-it using an adaptive Lagrangian controller that eliminates manual per-layer hyperparameter tuning.
A complete layer-by-layer SAE atlas for Gemma-4-E2B, trained and published live as each layer completes. Each SAE decomposes the residual stream activations at that layer into a sparse dictionary of 49,152 learned features.