Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Spring 2026 projects

Testing models’ self-forecasting abilities

Evaluations AI control Other

There are increasingly many tools for studying multi-turn agentic rollouts, like Petri, BrowserGym, and WebArena. Given the source code of a scenario, can the model predict what it will do when prompted within the scenario?

Your role

See proposal

Prerequisites

Comfortable using Cursor / Claude Code

Application question(s)

  1. Submit your portfolio (previous research, writing, Github repos, etc.)

  2. Why are you interested in this project?

  3. [optional] Read https://collisteru.substack.com/p/the-object-level-career. What are your object-level career goals?

About the mentor

Lydia Nottingham

Lydia Nottingham

CBAI, University of Oxford

View profile

Lydia researches LM behavior through CBAI to improve the predictability of frontier AI systems. She has worked on utility engineering with the Center for AI Safety and cooperative MARL at the Foerster Lab. Integrating tools from across AIS subfields, she pursues principled approaches to ML safety.

Similar projects