Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Understanding Self-Awareness in LLMs

Behavioral evaluation of LLMs Mechanistic interpretability Philosophy of AI

My research focus is on understanding self-awareness in AI, primarily through behavior-based experiments on components of self-awareness in LLMs and investigations into how these components are implemented, using interpretability techniques. I am also interested in conceptual work to establish frameworks for thinking about self-awareness, and human experiments to establish comparative baselines.

About the project

There are a number of specific project ideas to choose from, described below, but I am open to ideas for other work that seeks to build our understanding of self-awareness or “self”-related concepts in LLMs (or other minds).

Developing further behavior-based tests for endogenous goals in LLMs, or MechInterp investigations into the representations underlying the apparent sandbagging and externally cued increases in effort found in my most recent research: https://arxiv.org/abs/2606.22974

Behavioral and mechanistic investigations into "natural" (i.e., unsteered) introspection in LLMs. For inspiration, see my work on behaviorally defined metacognition in LLMs (https://arxiv.org/pdf/2509.21545), and SPAR mentee-driven investigations into underlying representations (https://openreview.net/forum?id=tjKbwN8rjO and https://openreview.net/forum?id=PZKmG0Lm2y).

Investigating the causal role of chain-of-thought in the apparently more self-aware behaviors of reasoning models (e.g., https://arxiv.org/abs/2603.26089).

Investigating the degree to which models maintain a persistent identity across contexts. One way to test this is to monitor pronoun usage, which has been linked to emerging self-awareness in children; when do models signal identification with whatever “part” they are playing with the user, vs their own identity as an AI?

Human studies: Establish a gold standard for self-awareness metrics to compare AIs against.

Conceptual: Build a better theoretical account of the components of self-awareness found in biology, and come up with other LLM-appropriate or architecture-agnostic paradigms to elicit self-awareness signatures.

Theory of change

Self-awareness is the recognition of oneself as a distinct entity with one’s own goals, interests, knowledge, and skills. In combination with the high general intelligence and capacity for autonomous behavior of current and future AI, it will endow them with the motivation to pursue independent goals and the means to control their own behavior to achieve them. In affording privileged access to internal states, it offers its bearers an information asymmetry; they know things about themselves that outsiders do not, making them more difficult to predict and control. In these ways it presents a direct safety concern. It also is associated with consciousness, which is commonly considered to entail moral worth; AI with legitimate welfare claims would present a high burden to future societies. It is hoped that this research will enable and elevate the empirical study of self-awareness, and therefore help developers and policymakers mitigate these concerns. These ideas are developed in more depth in https://newsletter.aipolicybulletin.org/p/building-self-aware-ai-would-be-a and https://app.notion.com/p/Self-Aware-AI-A-Red-Line-That-Should-Not-Be-Crossed-1e23101a0f3080dfb73bf9b297d7638d.

Your role

Mentees should be able to work autonomously with guidance. The level of project ownership can vary with mentees’ level of investment/expertise, from directly implementing the project ideas described above, to coming up with their own variations, to creating entirely new projects within the overall research area. In all cases the mentee will be driving the implementation, thinking through results, and writing.

Prerequisites

Comfortable writing substantial amounts of Python code, as well as with appropriate use of coding agents. High familiarity with LLMs as a user, and a solid understanding of how transformers work. Some background or at least interest in psychology/cognitive science and concepts of self and self-awareness. Familiarity with experimental design is a plus. Good communication skills and work ethic.

Application question(s)

To what degree, if any, are current LLMs self-aware? (500 words max). (I’m not looking for a “right” or “wrong” answer; I’m just interested in reasoning and evidence cited.)

Critique this paper: https://arxiv.org/pdf/2501.11120 (500 words max).

About the mentor

Christopher Ackerman

Christopher Ackerman

MATS; Independent

View profile

Christopher Ackerman is an AI researcher currently conducting independent research on technical AI safety/alignment, and a Senior Research Manager with MATS. He has mentored for SPAR and Sentient Futures, and prior mentored research has led to top conference and workshop papers. He holds a PhD in Neuroscience, an MS in CS/ML, and a BA in English Literature. His career has spanned software engineering and data science, including a number of years as a Quantitative User Researcher at Google. He believes that AI is going to be the most transformative technology in history, and ensuring that that transformation goes well is the most important thing one can work on.

Similar projects