Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Why does data attribution not work in realistic settings

Mechanistic interpretability Alignment

Data attribution has been shown to work in very simple and artificial settings. Recently it has been shown to not work that well on more realistic settings. Why is that?

About the project

Using data attribution to find training examples that lead to certain model behavior has been shown to work in simplistic scenarios (emergent misalignment). Recently, some exploratory work (https://www.alignmentforum.org/posts/aTybJ6CPQrxEY8rE2/data-filtering-works-a-lot-worse-than-you-would-expect) has shown that for other (realistic) traits, in a more realistic setting (SFT on real training documents), influence functions don't perform better than random. Reproducing this research is extremely important as it is one of the first negative results on more realistic settings, possibly undermining the possibility of using data attribution in its current stage, to understand the behaviors of models. Additionally I would like to extend these experiments to the DPO stage.

Theory of change

Data attribution is significantly under-explored technique in the AI safety community. Being able to tune the training data to best control model behavior is a very desirable technology and not only that, it might be possible to predict model behaviors/generalization properties before training them. Giving that this technology is still in its nascency, reproducing negative results is of utmost importance.

Your role

I expect mentees to run and write the code for the experiments. I will be working with mentees to come up with the experimental design, but I expect them to be able to think about interesting experiments to run

Prerequisites

Proficient at Python and Pytorch (or be able to prompt LLMs and evaluate the legitimacy of their outputs) Have done some finetuning experiments, as well as evaluated models (same LLM clause as above)

Application question(s)

What could be some things that went wrong in https://www.alignmentforum.org/posts/aTybJ6CPQrxEY8rE2/data-filtering-works-a-lot-worse-than-you-would-expect? (300 words)

How would you go on about reproducing these results, and are there other behaviors you think are more relevant?

About the mentor

Gonçalo Paulo

Gonçalo Paulo

EleutherAI

View profile

I'm an interpretability researcher at EleutherAI. I studied physics in university (bachelors, masters and PhD) before realizing that making models safer and understanding them is more interesting and important. I worked on automated interpretability using sparse autoencoders and am currently interested in data attribution and how we can predict the effect that data will have on model's behaviors.

Similar projects