Views
No views yet
| Llama Scope | Scaling Monosemanticity | GPT-4 SAE | Gemma Scope | |
|---|---|---|---|---|
| Models | Llama-3.1 8B (Open Source) | Claude-3.0 Sonnet (Proprietary) | GPT-4 (Proprietary) | Gemma-2 2B & 9B (Open Source) |
| SAE Training Data | SlimPajama | Proprietary | Proprietary | Proprietary, Sampled from Mesnard et al. (2024) |
| SAE Position (Layer) | Every Layer | The Middle Layer | 5/6 Late Layer | Every Layer |
| SAE Position (Site) | R, A, M, TC | R | R | R, A, M, TC |
| SAE Width (# Features) | 32K, 128K | 1M, 4M, 34M | 128K, 1M, 16M | 16K, 64K, 128K, 256K - 1M (Partial) |
| SAE Width (Expansion Factor) | 8x, 32x | Proprietary | Proprietary | 4.6x, 7.1x, 28.5x, 36.6x |
| Activation Function | TopK-ReLU | ReLU | TopK-ReLU | JumpReLU |
@article{he2024llamascope,
title={Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders},
author={He, Zhengfu and Shu, Wentao and Ge, Xuyang and Chen, Lingjie and Wang, Junxuan and Zhou, Yunhua and Liu, Frances and Guo, Qipeng and Huang, Xuanjing and Wu, Zuxuan and others},
journal={arXiv preprint arXiv:2410.20526},
year={2024}
}