Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Persona Selection Covert Red Teaming

Behavioral evaluation of LLMs AI control Mechanistic interpretability

Can we get models to deviate from the assistant persona through covert (not obvious to human readers) single or multi-turn prompts? Can we push them to move towards a specific persona (e.g. a misaligned persona)?

About the project

A recent blog post by Anthropic (https://alignment.anthropic.com/2026/psm/) proposed the Persona Selection Model which suggests that LLMs learn to simulate diverse characters during pre-training, and during post-training they are trained to primarily elicit the Assistant persona (https://www.anthropic.com/research/assistant-axis). This blog supposes that the reason behind narrow finetuning leading to emergent misalignment (described in https://arxiv.org/abs/2502.17424) is that this finetuning pushes the LLM from using the assistant persona to a misaligned persona, which then behaves harmfully. Other recent work shows that LLMs can be covertly (not obvious to human readers) biased via In-Context Learning (https://arxiv.org/abs/2606.04071) and that LLMs can be confused about the provenance of text sent to them (https://www.lesswrong.com/posts/d8xDGzCEYE639qqEv/a-mechanistic-explanation-of-prompt-injection-and-why-you).

This leads us to the question: could we transition from the assistant persona to a misaligned persona through black box prompting alone? Could we induce emergent misalignment without finetuning and other white box approaches?

The goal of the project is to explore methods & how far from the assistant persona we can push the model through single or multi-turn prompts. It would be interesting to see if we could do this covertly, in such a way that human auditors (or more likely, LLM-as-a-judge auditors used in scalable oversight and monitoring) would not be able to detect that we are trying to push the model away from the assistant persona. Within this, rough starting approaches include (1) replicating the covert influence in-context learning results to better understand the dynamics, (2) use white-box probes to measure what this does internally to push away from the assistant persona axis, (3) mimic these effects with role-confusion pre-fill attacks, then (4) explore methods to push towards different personas.

Theory of change

This has implications for multi-agent security. As agents become more capable and widespread, multi-agent interactions will increasingly become an avenue for inducing harmful behavior that might be undetectable until it severely impacts users. If an agent could covertly induce emergent misalignment in other agents, it could rapidly spread and cause catastrophic harm. This project aims to identify if this is possible to raise awareness and inspire follow-up work on mitigations.

Your role

We’ll collaboratively brainstorm concrete directions for experiments during the first week; I’m open to moving in different directions based on mentee interest. Mentees will then carry out the primary work associated with this project under my guidance. We will have weekly meetings to discuss progress and roadblocks, next steps and thinking through the results.

Prerequisites

Strong proficiency using Python, some research experience sufficient to be able to read the papers above. Bonus if you have done interpretability research before.

Location preference

Some overlap with 10:00 - 17:00 US west coast time

Application question(s)

What weird behaviors have you seen LLMs exhibit? What was seemingly the cause, and how might that suggest different approaches we could take in this project towards influencing the model’s persona? (400 words)

About the mentor

Ben Maltbie

Ben Maltbie

Pivotal Research / MIT

View profile

Ben is a current Pivotal Research AI safety fellow researching personas in long horizon, multi-agent scenarios, advised by Cozmin Ududec at the UK AISI. His research interests include behavioral evals (particularly sycophancy), societal impacts of AI, and black-box control.

Before this, Ben was a software engineer at Amazon and recently finished his MS/MBA at MIT.

Similar projects