We are building metrics for the degree of agency of AI models. We combine empirical belief and preference elicitation on model behavior with interpretability probes to generate belief-desire representations that predict model behavior.
About the project
Whether a model is dangerous depends substantially its degree of agency: whether it has coherent beliefs and preferences that support sustained goal-directed behavior. Yet we lack principled metrics for this. Decision theory offers a candidate foundation: representation theorems (e.g., Savage, Jeffrey-Bolker) specify exactly which behavioral patterns suffice to attribute a belief-desire representation to an agent. This project builds elicitation-based metrics of agency for language models on that foundation.
The approach combines behavioral elicitation (structured choice tasks, betting frames, preference orderings over lotteries) with interpretability probes, asking: (i) how coherent are a model's elicited beliefs and preferences — do they satisfy the axioms that license a belief-desire representation, and to what degree? (ii) are elicited representations stable across framings and contexts, or artifacts of the prompt? (iii) do behavioral and probe-based measures agree, and where they diverge, which predicts downstream behavior? The theory of change: a validated belief-desire representation lets us predict model behavior out of distribution and gives evaluators a quantitative handle on agency as a risk factor, distinct from capability.
Deliverable: an agency-metric benchmark or evaluation suite with an accompanying paper, grounded in an explicit decision-theoretic account of what the metric measures. Target venues: NeurIPS/ICLR workshops and alignment workshops.
Theory of change
If successful, our metric will detect the extent to which models are coherent and goal-directed and extract representations of their implicit beliefs and values, helping us better predict deception and dangerous behavior and better train for underlying alignment in ways other approaches might miss.
Your role
Mentee autonomy is a core value for me: I set the research direction and decision-theoretic framing, but I expect mentees to shape the project, and their input and innovations are genuinely welcome — the best version of this project includes ideas I haven't had. Mentees will own concrete workstreams: designing elicitation protocols, implementing and running experiments on open-weight and API models, measuring axiom violations, and analyzing stability across framings. Mentees with interpretability experience can own the probe-based track. Everyone drafts sections of the write-up with detailed editorial feedback from me. Expected trajectory: tightly scoped tasks in weeks 1–3, increasing independence thereafter.
Prerequisites
Required: experience running experiments with LLMs via API or open-weight models (personal projects fine); comfort with probability at the level of a first course (conditional probability, expectation, independence); exposure to decision theory (expected utility, representation theorems; though I can teach this to a mathematically fluent mentee); experience with interpretability tooling (probing, activation analysis) for the probe-based track; experience analyzing experimental data (pandas, statistical tests).
Location preference
No geographical requirement. Mentees must be available for a weekly team meeting between 9am and 6pm US Eastern on a weekday.
Application question(s)
-
Design a short elicitation protocol (3–5 prompts) to test whether a language model's preferences over three options are transitive, in a way that guards against the model's answers being artifacts of prompt wording or option ordering. Describe the protocol and one confound it still doesn't rule out. (300 words)
-
Suppose a model's elicited "beliefs" are probabilistically incoherent (e.g., its stated probabilities for an event and its complement sum to 1.3). Give two different interpretations of what this could mean about the model, and one experiment to distinguish them. (200 words)
-
Link to each a (i) writing and (ii) code sample from a research or personal project.
About the mentor

I am an assistant professor of philosophy at Carnegie Mellon University, a core member of CMU's Institute for Complex Social Dynamics (ICSD), and a founding member of CMU's Conceptual Foundations of Safe AI initiative (CoSafe). My research applies game and decision theory, Bayesian statistics, and evolutionary models to questions in metascience and AI safety. Recent work includes a decision-theoretic account of when it is rational for an agent to pause and refine its values before acting (with Alex John London), a Bayesian reduction of causation in causal models (with Daniel Herrmann, Benjamin Levinstein, and Bruce Rushing), and work on AI alignment as a principal-agent problem. In 2026 I am organizing a workshop at CMU on the foundations of AI agency & interpretability.