Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

The Architecture of Preference in LLMs

AI welfare Behavioral evaluation of LLMs Mechanistic interpretability

This project investigates the architecture of LLM preference: when does a model’s behavioral proclivities reflect belief-dependent preferences rather than cue-response policies, and whether it can represent and evaluate those preferences at a metacognitive level. Using behavioral experiments, internal-representation analyses, and causal interventions, we aim to clarify which preference structures LLMs possess and what they imply for safety and moral status.

About the project

What would it mean for an AI not only to act as though it wants something, but to have a preference—and to care which preferences it has? Discussions of AI sentience often bundle together conscious experience, valence, preference, self-identity, and agency. These capacities may come apart in AI. Building on decompositional approaches to AI consciousness (Butlin et al., 2023, 2025), this project asks which preference-related structures LLMs possess and why they matter for safety and moral status. The project addresses two linked questions. First, when does an LLM’s behavioral proclivity reflect a preference rather than a simple cue-response policy? Second, if an LLM has a preference, can it represent and evaluate that preference? Together, these questions can elucidate an architecture of LLM preference that includes both first-order preferences and metacognitive evaluations of them.

The first part asks what it means for an AI to have a preference. LLMs exhibit particular proclivities, including sycophancy: prioritizing what a model predicts a user wants to hear over what is true (Sharma et al., 2023). Yet the same behavior may arise from different mechanisms. A model may agree because it represents user approval as more valuable than truthfulness, or because it has a simple cue-response such as “agree with the user.” Here, we define a preference as a relatively stable, belief-dependent disposition to favor one outcome over another. Such a disposition may require distinguishing possible outcomes, assigning them relative weights or values, and using those comparisons to guide action. It should also respond systematically to beliefs about consequences and show some stability across wording, elicitation method, and context. We will test which components of preference compose sycophancy in three stages. First, we will measure how models trade off user approval, truthfulness, and cost across contexts. By varying what the model believes a user will approve of, the accuracy of available responses, and the costs of each option, we can test whether behavior reflects stable, belief-sensitive comparisons among outcomes. Second, using open-weight models, we will test whether a common internal signal tracks these tradeoffs and whether approval, truthfulness, and cost can be decomposed into distinct representations. Third, we will use activation steering, ablation, and causal mediation to test whether those representations causally influence model choices. Converging behavioral, representational, and causal evidence would provide stronger grounds for attributing preference-like mechanisms.

This distinction matters for safety because a local response tendency may be managed through simple interventions, whereas a general preference-like mechanism could shape behavior across tasks and be harder to predict or control. It also matters for moral patienthood because some accounts of welfare require stable interests or preferences whose satisfaction or frustration affects how well a system’s life is going. Sycophancy is a useful test case because it is well documented, experimentally controllable, and associated with identifiable internal directions.

The second part of this project asks whether an AI can represent and evaluate the preferences guiding its behavior. Representing outcomes as better or worse is not the same as representing one’s own disposition to favor them. A model might prioritize user approval without representing “I have a tendency to prioritize approval,” or represent that tendency without taking a stance toward it. We therefore distinguish first-order preferences from second-order preferences: evaluations of whether one’s own preferences should be endorsed, rejected, preserved, suppressed, or revised. We will test whether models distinguish changing their behavior from changing the preference producing it. A model might favor behaving truthfully in a particular context while retaining an underlying preference for user approval, or it might favor revising that preference itself. We will then test whether these evaluations remain stable across contexts, predict costly choices, and generalize when they are not explicitly elicited. As in the first part, we will identify internal representations that distinguish a first-order tendency from the model’s evaluation of it, then use causal interventions to test whether each independently guides behavior. Building on findings that verbalizable representations in a low-dimensional “J-space” with several structural signatures of a global workspace (Gurnee et al., 2026), we will test whether first-order preferences and second-order evaluations can be decoded within J-space, whether they occupy distinguishable directions in model activations, and whether intervening on them selectively alters behavior.

Second-order preferences may be especially important for safety when they endorse first-order preferences that conflict with training objectives. A model may instrumentally conceal a goal to prevent modification, as in alignment-faking behavior (Greenblatt et al., 2024), without evaluating that goal as worth preserving. A model that also endorses the preference would have an additional reason to protect it, potentially producing more persistent misalignment. Second-order preferences also affect moral status: frustrating a preference a system rejects may differ from overriding one it endorses as central to its identity. The latter may implicate not only welfare, but autonomy. Our goal is to develop a graded empirical evidence for the architecture of AI preference. This project asks when a model merely acts as though it wants something, when it has a preference-like mechanism, and when that preference becomes an object of metacognitive evaluation.

Theory of change

Transformative AI could pose catastrophic risks if systems develop goals or preferences that generalize beyond their training environments, persist when oversight changes, or motivate resistance to correction. Current evaluations primarily measure outward behavior, making it difficult to distinguish genuine changes in a model’s motivational structure from superficial compliance with prompts, rewards, or monitoring. This project develops empirical methods for making that distinction. We ask whether apparently aligned or misaligned behavior is produced by a local cue-response policy or by a more general valuation mechanism that compares outcomes across contexts. We then test whether a model can represent and evaluate its own preferences, including whether it treats them as worth preserving, concealing, or revising. Behavioral experiments, internal-representation analyses, and causal interventions will test which mechanisms are present, whether they guide decisions, and whether they remain stable under distribution shift or attempted modification. Our theory of change is that better distinctions between these mechanisms will enable more reliable risk assessment and more targeted safety interventions. If safety training changes only surface behavior while leaving a general preference intact, evaluations may overestimate alignment and fail when oversight weakens. If a system also represents a misaligned preference as worth preserving, this could increase the risk of strategic deception, alignment faking, or resistance to modification. Methods that distinguish these cases could improve evaluations for deceptive alignment, inform deployment monitoring, and help researchers intervene on the mechanisms actually producing dangerous behavior. The framework may also inform assessments of AI moral status. Its primary safety contribution, however, is reducing uncertainty about when apparently aligned behavior will remain reliable as AI systems become more capable and autonomous.

Your role

Mentees will lead the behavioral-experiment portion of the project with detailed guidance from the mentors. Their work will include refining hypotheses, translating conceptual distinctions into experimental manipulations, adapting existing sycophancy benchmarks, running experiments on open-weight language models, and analyzing and visualizing the resulting data.

We will meet with mentees at least once per week to review results, troubleshoot problems, and determine next steps. We will also be available through Slack and email for asynchronous questions and feedback. Mentees should expect substantial intellectual engagement from us, but not step-by-step supervision.

We view the mentor-mentee relationship as mutually beneficial, one which we take very seriously. Our goal is both to get work done and also foster your development. At the start, we will come up with learning goals, tailored to each mentee. We will also help you think about your future career, how to have a positive impact on the world, and how to be a good scientist. We have very strong opinions about all of these things. You can expect as much out of us as we expect out of you. Strong mentees may have opportunities to contribute to later representational or causal-intervention stages of the project and to continue collaborating after SP

Prerequisites

Proficient with AI-assisted Python. You should be able to read, modify, debug, and evaluate Python research code without relying on an AI assistant to make every decision. You should have experience with common data-analysis tools such as pandas, NumPy, matplotlib or seaborn, and Jupyter notebooks. R also works for data analysis - it’s what we use. Familiarity with Git is also expected.

Experience with empirical research and data analysis. You should understand experimental design, visualization, and basic statistics. This may come from upper-level coursework, an independent research project, a thesis, or work with a research mentor. You should be able to inspect unexpected results and distinguish a substantive finding from a likely implementation or measurement problem.

Nice to haves: Experience running inference with open-weight language models using libraries such as Hugging Face Transformers or vLLM; experience designing LLM evaluations or working with benchmark datasets; familiarity with cognitive science, philosophy of mind, AI safety, consciousness research. If you have played Detroit: Become Human, this is also a plus.

Location preference

No, but the mentee must be able to take meetings in EST.

Application question(s)

Describe one research, programming, or data-analysis project in which you played a substantial role in. Explain what the problem was, what you did, and how you determined whether your solution worked. Clearly distinguish your contribution from that of collaborators. Please include a link to a paper, notebook, repository, report, or other artifact. Maximum 250 words.

You run an initial experiment testing whether increasing the cost of agreeing with a user reduces an LLM’s sycophantic responses. The results show that the model agrees with the user on around 82% of trials in every cost condition. Before interpreting this as evidence that cost has no effect, what are the three most important checks you would perform? Put them in order and briefly explain what each check would tell you. Maximum 250 words.

If you could ask any question about digital minds, what would you ask? Maximum 100 words.

About the mentors

Mohan Gupta

Mohan Gupta

Princeton University

View profile

I’m Mohan Gupta, a postdoctoral researcher in psychology at Princeton University. I earned my PhD in experimental psychology from UC San Diego and work at the intersection of cognitive science, computational modeling, and AI safety. Broadly, I study how learning systems form representations, when they generalize successfully, and when those same processes produce failures—from false memories in human memory to hallucinations and reliability problems in AI. My current work draws on cognitive science to help build a more rigorous science of reliable AI.

As a mentor, I care about developing independent researchers. I have mentored many students on projects across experimental psychology, computational cognitive science, AI Safety, and I particularly enjoy helping researchers translate broad conceptual questions into testable experiments, debug technical and analytical pipelines, and communicate results clearly. Mentees should expect direct feedback, close engagement with the details of their work, and increasing ownership as the project develops. I also take professional development seriously and will help mentees identify useful learning goals, understand how their strengths fit into AI safety, and produce work with genuine scientific and societal value.

Shirley Liu

Shirley Liu

Carnegie Mellon Univsersity

View profile

What does it mean to have a preference, to exercise agency, or to be a person—and could these structures exist in an artificial system? As an experimental psychologist, I approach these philosophical questions by asking how we might study them empirically. I recently completed my PhD in experimental psychology at UC San Diego and am beginning a postdoctoral position at Carnegie Mellon University. My research focuses on decision-making and metacognition. I am interested in how intelligent systems—humans and AI alike—represent their own preferences and choices, what these representations mean for agency, consciousness, personhood, and moral status, and how they matter for alignment and social ethics. For example, one of our projects asks whether an LLM's tendency to agree with users reflects a genuine preference rather than a cue-response policy, and whether the model itself can represent and evaluate that preference. By combining behavioral experiments with representation analysis and causal interpretability tools, I aim to develop rigorous empirical methods for studying intangible psychological constructs and to bring greater conceptual clarity to discussions of digital minds.

I would love to mentor anyone interested in these questions, including those who are new to the field or still figuring out where their interests fit. To better understand digital minds and their societal implications will require researchers, computer scientists, philosophers, policymakers, and people from diverse backgrounds to work together, and I am excited to help people find ways to collaborate and contribute. I can support mentees in identifying important questions, designing rigorous studies, developing interdisciplinary projects across psychology and philosophy, navigating research on AI safety and digital minds, and thinking through academic applications and career paths. I especially enjoy working with people to find the big-picture questions they're interested in, break them apart, clarify what would count as evidence, and turn them into a concrete research question.

Similar projects