3,000 interpreted and verified SAE features across all 60 layers of Google's Gemma-4-31B-IT model.
For each of the 60 transformer layers in Gemma-4-31B, we trained a TopK-64 Sparse Autoencoder with 43,008 features (8x expansion from d_model=5376). We then selected the 50 most interesting features per layer using SIPIT (Sparse Input-Token Invertibility Probe) scores, interpreted them with two independent LLMs… See the full description on the dataset page:
https://huggingface.co/datasets/Adam1010/gemma-4-31b-sae-features.