Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Methods and measure for weight and representation based early detection of major changes during learning

Developmental interpretability Mechanistic interpretability

We will build theoretical models to study sudden changes in network internals during learning, deriving methods and measures from the theory to create a taxonomy of these changes and identify them early during training.

About the project

Neural networks, despite a smooth loss, undergo abrupt changes in their internal structure and behavior. Identifying these changes is important to detect when a network might have acquired a new capability or a different way of understanding the world. Behavior-based measures are insufficient for the task, since they require curated tasks and are too expensive to administer regularly through training. We aim to develop behavior-evaluation-free methods and measures to identify sudden changes in network internals based purely on weight structure and hidden representations (for non-curated inputs). The starting point will be my own ongoing theoretical work that studies representation formation, structure, and their timescales throughout training in solvable models.

Theory of change

As AIs saturate evaluation measures, we need to find ways to concretely track the development of AIs across training time and generations, in the hope that we can identify training time points or model generations that might have undergone a sudden break and could be substantially different from previous generations or checkpoints. Developing weight- and representation-based measures can serve as a first-pass early warning system to identify crucial time points that might mandate stopping and doing more thorough evaluations to understand the delta.

Your role

Mentees can either work on mathematical models where we can prove someting concrete about various measures and create a scientifically grounded taxonomy of various possible changes and appropriuates methods to identofy them.

Or they can work on scaling and applying methods already identified in my recent work to real world problem and validating the methods.

Prerequisites

  • Skilled in using python and knows how to steer and use LLMs for scientific projects
  • knows how to work with and manipulate internals of trained networks
  • understands linear algebra and probability proficiently

Application question(s)

Distinguish between lazy and feature learning regimes in mathematical models of neural networks. Discuss what you would expect to see in their weight structure and hidden representations.

About the mentor

Nischal Mainali

Nischal Mainali

Principles of Intelligence

View profile

I am a research scientist interested in theoretical understanding of AI systems and using that understanding for interpretability, i.e., theory-first interpretability. To that end, I am currently interested in exactly solvable models of learning dynamics and their applications, and mean-field theories of learning and representation formation. I would be interested in either pushing the theory frontier and developing basic understanding of safety-relevant phenomena (e.g., silent alignment, representation structure in mean-field networks, solvable models of superposition, etc.) or applying various ideas coming from theory for practical purposes (early detection of sudden changes in NNs during learning, new weight space/representation space interpretability tools, etc.). I am generally pretty open to exploration and will support mentee-led directions, but also have concrete problems and projects to work on.

Similar projects