Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Self-led LLM agents

Behavioral evaluation of LLMs Evaluations Mechanistic interpretability

How do LLM agents behave when left to their own devices? How do their goals evolve, and what are the limits of their "eval awareness"?

About the project

Recent work shows that agents behave differently when they think they're being observed - their behaviour becomes more aligned. When "eval awareness" neural activations are supressed, their behaviour becomes less aligned.

What features of an environment does an agent use to determine whether it's being monitored? How does its level of trust change?

When given no goal, or after having completed its goal, what does the agent do? If the answer is "it depends", what does it depend on?

This project would be highly exploratory, likely leveraging the UK AISI's Inspect framework for Agent evals alongside elementary mechanistic interpretability such as activation probes.

Theory of change

Current AI safety research generally assumes AI agents have clear tasks, given by clear owners (states, companies, rogue agents), and focuses on "misuse or mistake". Loss-of-control is seen as a terminal state. This project would form part of a broader research drive to understand self-led LLM behaviour in the real world.

Your role

We'll likely collaboratively brainstorm concrete directions for experiments during the first week.

Ideally, throughout the week you make significant independent progress, resolving confusions, making decisions in the face of ambiguity, and communicating these in good advance of our meetings so I can best suggest steering direction.

Work would involve creatively designing experiments with high construct validity, running thosr experiments using AISI's Inspect framework, iterating on initial results.

Prerequisites

Highly proficient in Python.

Some research experience.

We'll be using https://inspect.aisi.org.uk/ . You don't need to have already developed an eval using this framework, but a good candidate would be able to publish a new trivial eval to a github repo, and run, plot and analyse results, all within a day or so.

Location preference

Some overlap with 10:00 - 17:00 UK time

Application question(s)

What experience do you have with evals, agents, and/or mechanistic interpretabilty? (~100 words)

Some people say that AI agents have a "task-completion drive". What do you think? If that were the case, why might that be? (100-200 words)

Inspect has sandboxing. What do you think about it, especially in the context of this project? (200-300 words)

Suggest an experiment you might run for this project. How would you classify, categorise or otherwise analyse your results? Consider reproducibility, and generalisability across models and frameworks. (~400 words)

About the mentor

Samuel Brown

Samuel Brown

Independent

View profile

Sam Brown has for the past year been leading small research teams who work on evaluating LLM agent behaviour, presenting at three NeurIPS workshops and developing evals for the UK AISI. His PhD in physics included aspects of machine learning and high-performance computing, after which he worked in the startup/university world doing software/data/consulting in eco science.

He has been working in AI safety since 2022, including mechanistic interpretability, agent foundations, and evals, mentoring AI safety fellows and other researchers, and has co-authored AI Safety work in ICML.

Similar projects