Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Measuring adaptability of LLM-agents to predict their real-world performance

AI strategy Evaluations

LLM agents have a "jagged frontier": impressive feats but surprising weaknesses. We're mapping this frontier, guided by RL (stochasticity is hard), animal and infant cognition (object permanence is hard), and using this map to predict performance at tasks in fields like software engineering and cyber security.

About the project

Our team (including SPAR alumni) have established a proof-of-concept: "proxy" evals focusing on specific skills (e.g. ability to handle chains of many simple steps) can be used to predict performance at SWEBench.

We've taken this work to LLM-evals@NeurIPS : https://drive.google.com/file/d/1EDB1AG3RoMRqOpJoGjmrFpDZ7ytv6lNm/view?usp=sharing

We're now extending it as part of the UK AISI's Challenge Fund, focusing on two "precursor skills": ability to handle non-stationarity (e.g. gradual or step-change in an environment) and stochasticity (e.g. noisy environment). See more detail here: https://docs.google.com/document/d/1ALFPmaS5ji7kBRIqA2fnXS1ovYaZYLwhjGatylx1F1I/edit?usp=sharing

As well as developing "proxy evals" for our two priority precurisor skills, using the Inspect Framework, you could look into other precursors you think are important (e.g. object permanence), or into the other "stretch goals / side-projects" listed in the linked Google Doc, e.g. writing "cheating metrics" (how often do agents reward-hack our evals?)

Theory of change

During pre-deployment testing of frontier models, UK AISI are often constrained by resource and time. AISI are providing us with CyBench results which we hope to be able to predict using our "proxies" approach - fast and cheap prediction could help efficiently target resources during pre-deployment testing.

Your role

For this project, the mentee could play a number of roles, according to their preference.

To work alongside the core team, a mentee would be taking considerable direction, following a clear path and using established methods.

Alternatively, mentees could explore complementary approached (e.g. the "stretch goals / side-projects" section of the GDoc linked above. In this case, they will have day-to-day autonomy in the research, carrying out experiments and taking initiative in exploring promising avenues. I expect to be most closely involved at the very start of the project during on-boarding (introducing tools, suggesting initial experiments), and towards the end during write-up, and expect to provide detailed guidance throughout.

It would be wonderful for results to be published in a workshop paper (side project) or conference paper (joining existing team). For this to happen, results should be rigorous. If mentees preempt ways in which initial experiments could be flawed, they can address these before our meetings to accelerate progress.

Prerequisites

We use use https://inspect.aisi.org.uk/ . You don't need to have already developed an eval using this framework, but also a good candidate would be able to publish a new trivial eval to a github repo within a day or so.

If you want to run more exploratory / side-project work, some research experience is necessary.

Location preference

Ability to overlap 1 hour 10am-5pm UK time

Application question(s)

Please provide a critique of the following paper (your own words, no LLM assistance please, 400 words): https://drive.google.com/file/d/1EDB1AG3RoMRqOpJoGjmrFpDZ7ytv6lNm/view?usp=sharing

Apart from those "precursor capabilities" mentioned above (ability to handle non-stationarity / stochasticity, object permanence), what other precursor capabilities might bottleneck agent performance at meaningful tasks?

About the mentor

Samuel Brown

Samuel Brown

Independent

View profile

Sam Brown has for the past year been leading small research teams who work on evaluating LLM agent behaviour, presenting at three NeurIPS workshops and developing evals for the UK AISI. His PhD in physics included aspects of machine learning and high-performance computing, after which he worked in the startup/university world doing software/data/consulting in eco science.

He has been working in AI safety since 2022, including mechanistic interpretability, agent foundations, and evals, mentoring AI safety fellows and other researchers, and has co-authored AI Safety work in ICML.

Similar projects