Current monitoring evaluations do not fully capture realistic internal deployments, where adversaries can adapt to defenses, exploit long-horizon interactions, and leverage environmental state. This project will focus on: (1) developing red-teaming environments where attacks emerge and co-evolve with monitors rather than being enumerated in advance, and (2) systematically evaluating which monitoring strategies and oversight levels remain robust across different threat models and deployment constraints.
About the project
Please visit this link for the project description.
Theory of change
As frontier AI systems become more capable and agentic, monitoring is likely to become an increasingly important component of practical AI control. However, current control evaluations largely measure performance against static benchmarks and dataset that do not reflect deployment. By developing more realistic red teaming environments and studying monitoring under adaptive adversaries, this work aims to identify failure modes before they occur in real systems and provide more reliable evidence about which monitoring and oversight protocols remain effective as capabilities improve.
Your role
- Own a concrete workstream within the project, depending on interest and fit. E.g., designing and building red-teaming environments (long-horizon, persistent-state, or multi-agent settings), implementing attack policies and automated red-teamers, or running and analyzing monitoring/oversight evaluations across control protocols.
- Contribute to the research direction itself: mentees who develop their own hypotheses within the agenda (e.g., a new attack class or monitoring strategy to test) will be supported in pursuing them.
- Co-author the resulting paper/conference submission, with authorship reflecting contribution; a mentee who leads a major component would be positioned as a lead or joint-first author on that piece of the work.
Prerequisites
- Good software engineering skills - You can write high-quality, legible code quickly (skilfully using AI), navigate large codebases you didn't write, and debug problems quickly and thoroughly. You have good judgment about when to write stable shared code vs throwaway code, and when to invest in a refactor, tooling or faster iteration loops. You check and understand your AI’s work.
- Technical skills: Python + libraries + Inspect AI + building and debugging evals.
- Hands-on experience running experiments with LLM APIs or agent scaffolds .
- Read and understood the core AI control framing - at minimum Greenblatt et al., “AI Control: Improving Safety Despite Intentional Subversion.”
- Able to commit minimum 10 hours/week for the duration of the program.
Nice-to-have (not required):
- Comfortable reasoning about classifiers, ROC curves, false positive rates, and base rates (much of this project is about measuring monitor performance correctly).
- Familiarity with control/monitoring evaluation work such as SHADE-Arena, Ctrl-Z, or the AI control safety case sketch.
- Experience with red-teaming, security (e.g., ATT&CK-style attacker modeling), multi-agent systems, game-theory or evaluaitons.
- Experience designing sandboxed or agentic evaluation environments.
Location preference
No preference. But
Application question(s)
- Our first goal is a realistic control setting in which attacks emerge and adapt against a monitor, rather than being written in advance. Assume no constraints, and sketch a concrete v1 of such an environment. Be specific about (a) the agent's legitimate task, (b) what counts as a successful attack, (c) how you would know, after the fact, whether an attack succeeded or the monitor missed one, and (d) why this design works, that is, why attacks would genuinely emerge and adapt against the monitor rather than staying fixed. Then name the single hardest design problem in your sketch and how you would handle it. (approx. 400 words)
- Briefly describe a stealthy or long-horizon attack strategy an AI agent could use to evade a transcript-based monitor, and why per-action monitoring might miss it (approx. 250 words).
- [Optional] What makes you a good fit for this particular stream? This could e.g. include a) previous research experience; b) experience with control or AI Safety; c) generally good SWE skills; d) anything else that you think indicates a strong fit. No need to be modest, feel free to show off. (approx. 250 words)
About the mentors

Sree is a MATS 10.0 scholar working with the UK AISI red team on AI control and loss-of-control risks. He previously worked with Redwood Research during the Pivotal Fellowship, also on high-stakes AI control. His main focus is high-stakes control research, and his goal is to study and design control protocols and monitoring techniques that are effective under real-world deployment conditions.

I'm a MATS scholar under Jacob Merizian (CAISI), working on red-teaming multi-agent capability, building robust monitorability and safeguards for agentic LLM systems. I'm also a Technical Member at EuroSafeAI, working on various projects to improve the safe coordination of advanced AI systems with critical human infrastructure.
My previous work included investigating scheming behavior in strategic multi-agent LLM interactions, LLM psychological phenomenas, and characterizing game-theoretic interactions in LLMs. I'm happy to support research in multi-agent AI control, as I think this is an incredibly impactful research agenda.