There are increasingly many tools for studying multi-turn agentic rollouts, like Petri, BrowserGym, and WebArena. Given the source code of a scenario, can the model predict what it will do when prompted within the scenario?
About the project
Theory of change
Your role
See proposal
Prerequisites
Comfortable using Cursor / Claude Code
Application question(s)
-
Submit your portfolio (previous research, writing, Github repos, etc.)
-
Why are you interested in this project?
-
[optional] Read https://collisteru.substack.com/p/the-object-level-career. What are your object-level career goals?
About the mentor

Lydia researches LM behavior through CBAI to improve the predictability of frontier AI systems. She has worked on utility engineering with the Center for AI Safety and cooperative MARL at the Foerster Lab. Integrating tools from across AIS subfields, she pursues principled approaches to ML safety.