We will conduct low-level mech interp research to better understand 1-layer neural networks.
About the project
Right now, we don't know how even the smallest of models work. Yet, a lot of research has focused on scaling existing techniques to larger models, without knowing whether they are truly finding the true underlying features and computations in networks. We will start over by looking at simple 1-layer language models to try to develop a more robust understanding of these networks.
We will be flexible with the techniques that we apply - this will be a mix of existing methods such as parameter decomposition, SAEs, and transcoders as well as developing new methods.
Here is a flavor of some of my existing work decompiling a 1-layer (non-language) model: https://www.lesswrong.com/posts/vGCWzxP8ccAfqsrS3/thoughts-about-the-mechanistic-interpretability-challenge-2
Theory of change
AI remains a black box and mech interp has not yet found the true underlying features and basis for computation for even the smallest of networks. Understanding how these small networks work will pave the way towards understanding even larger networks.
Your role
Mentees will conduct low-level mech interp research. This will involve writing code and deep thinking about mechanisms.
Mentees will meet with me about an hour a week and can also ask me questions as they arise.
Mentees will be responsible for writing their own code and be expected to debug their coding issues.
Prerequisites
Experience with Python and deep learning.
No other concurrent research projects.
Expect to be looking for a full-time AI safety/research role in 3-9 months.
USA-based.
Location preference
USA
Application question(s)
In my mind - mech interp is one giant puzzle - What are some of your favorite puzzles to think about?
What's something interesting about puzzles you've thought about?
What other commitments do you have this winter/spring?
Do you have any existing mech interp projects that you can share?
Which other mentors are you applying to?
About the mentor

Hi - I'm Rick Goldstein - I'm an independent mech interp researcher. I'm most passionate about digging into small models and trying to come up with a sufficiently complete explanation of the network's algorithm so we can handcraft a network to perform the task. Here's one example of this: https://www.lesswrong.com/posts/vGCWzxP8ccAfqsrS3/thoughts-about-the-mechanistic-interpretability-challenge-2
Currently, I've been thinking a lot about interpreting 1-layer language models which remains a challenging task and interpreting models trained to control robots.
Other background - Left my job in March 2023 to do interp full-time. Completed MATS in Winter 2024. Former SWE at Waymo PhD in Robotics from Carnegie Mellon University
I'm excited to see help people enter the field of mech interp.