This repository contains model checkpoints for Sparse Autoencoders (SAEs), as described in the paper
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability.
These models can be used for feature extraction.
Project page:
https://saebench.xyz.
For code, please see the
SAEBench repository.