SAE Lens
SAELens exists to help researchers:
- Train sparse autoencoders.
- Analyse sparse autoencoders / research mechanistic interpretability.
- Generate insights which make it easier to create safe and aligned AI systems.
HookedSAETransformer, SAEs can be used with Hugging Face Transformers, NNsight, or any other framework by extracting activations and passing them to the SAE's encode() and decode() methods.
Please refer to the documentation for information on how to:
- Download and Analyse pre-trained sparse autoencoders.
- Train your own sparse autoencoders.
- Generate feature dashboards with the SAE-Vis Library.
This library is maintained by Joseph Bloom, Curt Tigges, Anthony Duong and David Chanin.
Loading Pre-trained SAEs.
Pre-trained SAEs for various models can be imported via SAE Lens. See this page for a list of all SAEs.
Migrating to SAELens v6
The new v6 update is a major refactor to SAELens and changes the way training code is structured. Check out the migration guide for more details.
Tutorials
Join the Slack!
Feel free to join the Open Source Mechanistic Interpretability Slack for support!
Other SAE Projects
- dictionary-learning: An SAE training library that focuses on having hackable code.
- Sparsify: A lean SAE training library focused on TopK SAEs.
- Overcomplete: SAE training library focused on vision models.
- SAE-Vis: A library for visualizing SAE features, works with SAELens.
- SAEBench: A suite of LLM SAE benchmarks, works with SAELens.
Citation
Please cite the package as follows:
@misc{bloom2024saetrainingcodebase,
title = {SAELens},
author = {Bloom, Joseph and Tigges, Curt and Duong, Anthony and Chanin, David},
year = {2024},
howpublished = {\url{https://github.com/decoderesearch/SAELens}},
}