Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Exploring neuron monosemanticity

Mechanistic interpretability Developmental interpretability

Tracking the development of interpretable directions read by neurons through training. Understanding what incentives if any exist for MLP neurons to be monosemantic or for superposition to be contained to individual experts in MoE models.

About the project

  1. Understanding how MLP neurons learn their interpretable directions: literally looking at neuron activation dashboards through training and seeing if they become more or less monosemantic, and what data causes neurons to acquire their meaning; figuring out whether neurons which are interpretable "evolved" for a specific role. Speculative but "big if true".
  2. Understanding incentives MLP neurons have to be or not be monosemantic. Some toy model work with toy models of computation in superposition and teacher/student settings. MoE neurons are of special interest; can we prove features only interfere within specific experts and don't rely on other experts? What would it mean to prove this? What interpretability methods can we propose which resolve this superposition in a principled way?
  3. Finding specific cases of superposition between many neurons in MLP layers, finding ways to prove superposition exists and serves a specific role (as opposed to it just not being the case that there is a monosemantic neuron). Most speculative of the 3.

Theory of change

Labs currently use tools based on dictionary learning for monitoring; it seems plausible that these are actually not necessary and there are cheaper ways to resolve superposition, achieving this seems like it would save resources and make the methods the labs are using more well-motivated / reliable. More generally, this could help update the field towards / away from the feature hypothesis, which has kind of been in a stasis for about two years, and find better grounding for primitives we interpret models with.

Your role

I expect mentees to be self-guided and have very promising and clean ideas and be able to continue finding interesting things; that is my current bottleneck. Conditional on this, I will try to be as helpful as I can, provide frequent feedback, try to resolve bottlenecks, and help run experiments / write things up. We would essentially be more like colleagues than mentors/mentees.

Prerequisites

I am mostly selecting for research taste, i.e. ability to tell good research / methodology from bad and stay focused on promising directions. I.e.:

Good understanding of the transformer architecture, MLP neurons, and literature on SAEs and transcoders. Experience Claude coding.

Please only apply if you think you have a good idea that could help resolve one of the questions raised in the project description.

Application question(s)

All optional. Imagine you are already in the project and are talking about research you're planning to do. Please try to be laconic, but rambling is fine if it's coherent and not confused. Try to be well-calibrated.

  1. Propose an idea you would work on in this project.
  2. Do you think neurons in large language models are monosemantic? Provide evidence for / against
  3. What reasons would there be for neurons to be / not be monosemantic? How is the picture different for mixture of experts models and why?
  4. Do SAEs find "ground truth" features, a set of directions the model actually uses? Provide evidence for / against
  5. (extra optional) What do you think about the claim parameter decomposition methods are more well-motivated than SAEs because their number of alive components caps out with longer training

About the mentor

Stepan Shabalin

Stepan Shabalin

independent

View profile

Currently at Georgia Tech doing independent research on model diffing. Previously MATS 6.0 with Neel Nanda, intern at Eleuther, Anthropic Fellow. Interpretability interests: dictionary learning, neuron-based interpretability, metamodels, diffusion, characterizing superposition, simulators.

Similar projects