Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Alignment without Personas

Alignment Philosophy of AI

Alignment techniques that rely on 'personas' will fail as the AIs become more powerful. This project aims to make progress on the “hard problem of alignment” by providing a framework within which to measure how persona-dependent an alignment technique is, mapping and demonstrating the limits of persona-based alignment and of alignment techniques that make use of personas, and increasing awareness of the limitations of persona-based alignment techniques.

About the project

The following is a snippet taken from the project doc.

'Personas' are “the human-like characters appearing in text” that LLMs learn during pretraining Anthropic 2026. Much of AI Alignment is currently persona selection. Techniques such as prompting, fine-tuning, and distillation from conditioned models all require that the LLM learned moral human-like characters in its pretraining data. There are many reasons to expect persona-based alignment to fail. The following is a non-exhaustive list. The project will work to extend the list.

  • There will be no good role models for an aligned superintelligence to copy because the pretraining data has no aligned superintelligences.
  • Long-horizon outcome-based RL will push the model to have instrumental convergent goals
  • Even if the model did try to predict what a moral human would do in its place, there’s a pretty good chance that that a good human’s values would make for poor ASI values
  • The science of generalization with which we would be able to predict and manage inner misalignment is still a mystery Many alignment techniques work by pointing to a region in persona space for the model to mimic, some don't rely entirely on personas but are made more effective by the language models' prior, and some classic approaches do not rely on personas at all. The Alignment community needs to develop more production-ready techniques that do not rely on the AI behaving as a character from pretraining.

The main steps on the Roadmap (fleshed out in the project doc) are Gathering Perspectives on persona-reliance, Mapping AI Alignment, Testing How Well Alignment Techniques Work on a Model Organism of Persona-Absence, and Demonstrating Limitations of Persona-Based Alignment. We can prioritize these out-of-order, and the one I am most excited for is the Model Organism of Persona-Absence.

Theory of change

This project would help researchers see, understand, and track the limitations of current alignment methods and predict which future alignment methods are promising for out-of-distribution superintelligences. It would advance AI Safety to a great extent if we can create a quality metric for AI alignment techniques which the community can hillclimb.

Your role

Mentees will have a large amount of autonomy. My project proposal indicates a broad area of work that I want to see done, but I also expect mentees to have their own approach to the topic.

Prerequisites

Has engaged with AI Alignment (eg BlueDot course, a university AI safety fellowship, or equivalent). Has read at least one book on the alignment problem.

Application question(s)

Share a link to something cool you coded up with an AI coding agent which required a lot of guidance from you (at least 5 hours of human time).

Read one of the major works on personas. Write about details/considerations you find which are relevant for the Alignment Without Personas project. You get points for identifying difficulties, and even more points for figuring out ways we can get past those difficulties. (max 300 words, and please don't spend more than 1.5 hours on this):

(Choose one of the following questions)

  • What's a classic (pre-2021) AI alignment idea which deserves more attention in your opinion? Why? (250 words)
  • What's an AI Alignment technique you expect to work well on LLMs even without aligned personas for the LLM to mimic? Why do you think that? (250 words)

About the mentor

Matthew Khoriaty

Matthew Khoriaty

Pivotal AI Safety Research Fellowship

View profile

Hi! Currently I'm doing AI Control research at the Pivotal Fellowship with Redwood Research. Before this, I was the president and founder of Northwestern's AI Safety and Governance group, and at the ERA Fellowship, I developed a way to evaluate the diversity of a language model's output. In my thinking, I'm guided by classic LessWrong ideas surrounding questions like Instrumental Convergence and the Fragility of Value. Lately, many AI alignment researchers are thinking in terms of personas; I think that personas are obscuring a lack of progress on the alignment problem and we need to reconsider our approaches. Heads up: this would be my first time mentoring a research project. Although I think I will be good at it, it is only fair I inform you.

Similar projects