We build model organisms to study behaviors of interest (sycophancy, eval gaming/awareness, sandbagging, etc.) and test mitigations. But it’s unclear to what extent model organisms are realistic testbeds. We want to explore to what extent intervention success on model organisms is predictive of intervention success on base models.
About the project
We want to study how interventions on model organisms transfer to naturally emerging behaviors.
For each behavior of interest we want to study (sycophancy, eval gaming/awareness, sandbagging, etc.) we’ll find (i) a model that naturally exhibit the behavior, and (ii) a model organism specifically trained to exhibit it (e.g. a model organism of sycophancy).
We’ll choose a set of supposedly mitigating interventions (prompting, synthetic document fine-tuning, consistency training, white-box interventions, etc.) and for each of them:
- Measure how the intervention affects the model organism
- Use that to make a prediction on how it should affect the naturally misaligned model
- Measure how the intervention affects the naturally misaligned model
- Score our predictions against the actual results
Theory of change
Testing mitigations is one of the primary reasons for developing model organisms. For the most catastrophic forms of misalignment, we want to develop and test mitigations before the misalignment arises naturally, and are therefore restricted to model organisms as testbeds. However, model organisms are only useful here insofar as their response to the tested mitigations are predictive of naturally misaligned model responses to these mitigations. To ensure this, we need a science of model organisms. This project works towards such a science of model organisms by comparing how different model organisms respond to common mitigations with naturally misaligned models.
Your role
Mentees will be responsible for project execution (writing code, running experiments, analyzing results, writing down results, etc.). We’ll help with setting the research direction, and choosing which experiments to run.
Prerequisites
- Proficiency in Python
- Familiarity with AI safety fundamentals
- [Optional but encouraged] Experience with training, finetuning and evaluating LLMs
Application question(s)
- What’s a technical achievement you’re proud of? Please share links to any relevant public information if applicable. (<5 sentences)
- Can you describe a paper you’re excited about and say why it’s exciting? (<5 sentences)
- Why are you interested in this project? And what is your motivation for participating in SPAR? (<5 sentences)
About the mentors

Jeanne is a PhD student in the AI Safety & Alignment group at Max Planck Institute Tübingen.

Sohaib Imran
Independent
I am an independent researcher funded by Coefficient Giving. Previously, I was a Senior Research Associate and PhD student at Lancaster University.
I have primarily worked on LLM situational awareness. My earlier work concerned understanding internal reasoning in LLMs, including the potential of abductive out-of-context reasoning (https://arxiv.org/abs/2508.00741) and Bayesian updating (https://arxiv.org/abs/2507.17951). More recently, I've been working on a consistency training method to mitigate evaluation gaming (https://arxiv.org/abs/2606.02211). I have also worked on understanding self-awareness in LLMs (to be published). Other than that, I spend a lot of time thinking about deceptive alignment and scheming concerns.