This project aims to understand what personas are: in particular, we'd like to understand to what extent they're continuous.
About the project
We would like to know whether personas are best thought of as discrete (natural kinds), or whether they are highly continuous. We might think of personas as existing in a "persona landscape" (cf Assistant Axis and Persona Selection Model).
There are two main variables for surveying this landscape: the post-training and the prompt. My intuition is that post-training terraforms the landscape (making the Assistant basin larger and deeper; combining other basins into it), whereas prompting determines where you are within the landscape. (Post-training with the chat template in particular would strongly determine your “degrees of freedom” when accessing the landscape: there may be craters that you can’t access with the post-trained chat template but which nevertheless exist.)
We will begin by parametric variation of the prompt, aided by mechanistic interpretability, in order to search the landscape. The prompt variations can be guided by techniques from automated red-teaming, combined with following persona gradients found via mech interp. We’ll need some way to measure where in persona-space we are; this can also be found by mech interp (in which case we’d want to limit the mech interp used in finding the prompt variations in the first place, of course) or simply by LLM judges. If we believe we have found robust techniques for eliciting personas via this automated search of prompts, we can validate them by developing tools to measure the stability and coherence of the personas, perhaps by measuring perplexity.
Specifically, this might look like: Develop metrics for measuring persona elicitation, validated on manually verified transcripts: this could include projections onto the Assistant Axis (and other methods described in that paper to elicit personas); SAEs; LLM judges; along with others we’d discover via exploration, such as non-Assistant Axis persona steering vectors. Begin with a single prompt, and use the automated jailbreaking/red-teaming techniques described in the paper linked above to travel along the persona subspace. Characterize how robust initial persona seedings are to such variation: does it take a lot to get out of the Assistant, but not out of other personas? Can we find a reliable way of getting the model fully out of the Assistant basin without training it further: very out-of-distribution prompts, steering vectors, initialization of the residual stream to random values, etc?
This is the core research program: how resilient is the persona to perturbation? This will allow us to understand how large the persona is.
After describing the default persona landscape, the next question is where and why it emerges. First, we’d repeat the experiments on the base model: which personas does the base model have strong priors on? Second, we’d causally intervene on the post-training pipeline at various levels: how robust is the final Assistant to changes in the post-training pipeline? Can we use the tools we developed in the previous step — which will have hopefully yielded techniques to measure the stability of a persona — to make a maximally stable persona, which would have consequences not only for welfare but also for safety?
Theory of change
Personas are a possible explanation for misaligned behavior in the models: we might think of jailbreaks as a way of getting the model out of the default Assistant persona, or we might think that the Assistant persona is not sufficiently aligned. Characterizing how personas work is essential to understanding why they fail.
Your role
You'll have a lot of autonomy, and unfortunately, I won't be able to be an IC. You will be primarily driving the project and doing all experiments, though I'll be guiding you throughout, similar to most mentor-mentee relationships in the space.
Prerequisites
- Experience with empirical ML research, particularly mech interp
- Familiarity with persona literature
- Familiarity with AI welfare literature
- Philosophy background/experience is a plus!
Application question(s)
- Please describe your model of what a persona is, in as much mathematical detail or formality as possible (400 words).
- Please describe with as much detail as possible the first two experiments you would run, and what you would do next depending on the result (400 words)
About the mentors

My research focuses on AI welfare and mechanistic interpretability. I did the FIG Fellowship with Eleos where I worked on personas, the Anthropic Fellowship where I worked on mech interp + AI welfare, and the Cambridge Digital Minds Fellowship. I'm currently a PhD student at NYU, and did my undergrad in computer science and philosophy at Pomona College. I used to be a software engineer before pivoting to AI safety. A representative example of my work is a paper at functionalwelfare.com.
In my free time I like reading (especially poetry), writing, bouldering, and recently got into gunpla.

Ariana is a co-developer of recontextualization and has broad interests in shaping model generalization and model psychology. She is currently a researcher in the Anthropic Fellows Program, and recently worked on mitigations for reward hacking and misgeneralization as a MATS fellow with Alex Turner and Alex Cloud. She is also a SPAR alum. In the past, she has worked on training defenses against emergent misalignment and computer vision research.