Profile
Back to NewsBack
GitHub Trending 21 min
Reader Mode
AI-in-Transportation-Lab/awesome-mechanistic-interpretability: A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on Mechanistic Interpretability, a growing subfield i

AI-in-Transportation-Lab/awesome-mechanistic-interpretability: A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on Mechanistic Interpretability, a growing subfield i

7 hours ago

Awesome Mechanistic Interpretability

!Awesome License</a> !GitHub Contributors !GitHub Last Commit GitHub Stars</a> !GitHub Forks

A carefully curated collection of high-quality libraries, projects, tutorials, research papers, and other essential resources focused on Mechanistic Interpretability, a growing subfield in machine learning interpretability research that aims to reverse-engineer neural networks into understandable computational components. This repository serves as a comprehensive and well-organized knowledge base for researchers, engineers, and enthusiasts working to uncover the inner workings of modern AI systems, particularly large language models (LLMs).

To ensure that the community stays updated on the latest developments, our repository is automatically updated with recent mechanistic interpretability papers from arXiv. This ensures timely access to new techniques, discoveries, and frameworks that are shaping the future of model transparency and alignment.

[!NOTE]
📢 Announcement: Our paper from AIT Lab is now available on ACM CSUR!
Title: Bridging the Black Box: A Survey on Mechanistic Interpretability in AI
If you find this paper interesting, please consider citing our work. Thank you for your support!
@article{somvanshi2025bridging,
  title={Bridging the Black Box: A Survey on Mechanistic Interpretability in AI},
  author={Somvanshi, Shriyank and Islam, Md Monzurul and Rafe, Amir and Tusti, Anannya Ghosh and Chakraborty, Arka and Baitullah, Anika and Chowdhury, Tausif Islam and Alnawmasi, Nawaf and Dutta, Anandi and Das, Subasish},
  journal={Available at SSRN 5345552},
  year={2025}
}

Whether you are investigating the circuits behind in-context learning, decoding attention heads in transformers, or exploring interpretability tools like activation patching and causal tracing, this collection serves as a centralized hub for everything related to Mechanistic Interpretability — enriched by original peer-reviewed contributions and hands-on research from the broader interpretability community.

Updates

  • [Feb 04, 2026]: Our paper has been accepted at ACM Computing Surveys 🎉!
  • [Jul 24, 2025]: Preprint is now available in SSRN.

Last Updated

September 15, 2026 at 03:00:24 AM UTC

Theorem

Papers (1189)

Chat with me