Build the groundwork for a theory of what can and cannot be learned about agents. E.g. for a specific agent model that invokes goals, beliefs, intents etc, when can these endogernous variables be learned from behavioural experiments, and under what assumptions. Use this to derive fundamental identifiability results - in the first instance, to prove that the beliefs and goals of expected utility maximizers can be learned under weak assumptions, and derive scalable algorithms for doing so.
About the project
Everything that can be learned about an agent (including from activations) is ultimately grounded in behaviour. Endogenous agent properties (such as beliefs, goals, intents) are only valuable insofar as they allow us to predict agent behaviour. For example, training a deception probe requires labels for deception, which ultimately must come from behavioural experiments (or suffer confounding). And a deception probe is only useful insofar as it predicts if agents will behave deceptively or not, which in turn requires a behavioural definition. Despite this, the problem of determining if and how agent properties can be identified from experimental data is largely unexplored.
For example, it is often stated that the reward function in IRL is unidentifiable, but this is not true. It has been shown that if you consider the full causal treatment, e.g. you allow the experimenter to intervene on the environment and observe how the agent responds, you can recover its reward function [1].
How we will make progress on this problem in 3 months? The proposed project is a fairly self-contained identifiability result, which is unproven but I am confident is provable with a few months of effort. Starting with a general formulation of the agent property identification problem, we will build upon previous results (e.g. [2]) to prove that under weak assumptions both the goal and beliefs of a regret-bounded agent can be recovered from their behaviour alone. We will then explore relaxations of the requisite assumptions, and scalable algorithms for recovering goals and beliefs from behaviour, and demonstrate the efficacy compared to inverse planning and inverse RL.
[1] https://arxiv.org/abs/1601.06569 [2] https://arxiv.org/abs/2402.10877
Theory of change
Scalable algorithms for recovering the beliefs and goals of agents from behaviour, under testable assumptions, will hopefully be useful for any work looking to measure or describe goal-directed behaviour in AI systems.
Your role
High autonomy, mentees will be doing creative theorem proving, and (if successful) coming up with their own directions for expanding upon initial results. Work will mostly be pen-and-paper, with perhaps some empirical work towards the end if desired.
Prerequisites
Must be experienced with writing mathematical proofs, ideally having authored work in this area.
Location preference
No preference
Application question(s)
Please provide a link to relevant prior work
And / or a critique of papers [1] and [2] in the proposal.
About the mentor

Jonathan Richens
Google Deepmind
I'm a senior RS at google deepmind, working on agent foundations (theory), causality, and their application to alignment and safety. Before this I worked in quantum foundations and thermodynamics. More recently, Ive been working on selection theorems - kind of like representation theorems (Savage, Wald, etc), but instead of determining what agent models can be `fit' to behaviour, selection theorems try to determine what properties agents must have to achieve certain capabilities. Im most interested in attacking foundational questions about agency from novel angles, using simple tools and new ideas. This work is premised on the conviction that there are still pretty fundamental results laying around waiting to be found, without too much effort. If you enjoy proving theorems and doing a lot of highschool maths for long periods of time, come join my project!