Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

AI auditing under strategic attack selection

AI control Scalable oversight Evaluations

We will build and run a small experimental platform for AI auditing games in a synthetic transcript setting. The project will develop a toy environment, baseline attacker/defender policies, and evaluation code to test how different auditing strategies perform under limited budget and adaptive attack selection.

About the project

This project aims to make progress toward practical AI auditing by building a runnable simulation environment for strategic transcript review, rather than focusing primarily on abstract game theory. The core question is: when an auditor has limited budget and must inspect weak, noisy signals, what auditing policy is most effective against an adaptive system that can choose how to hide harmful behaviour?

This project will have different workstreams depending on mentee interest. The main deliverable will be an open-source toy framework that lets us run repeated auditing games with different assumptions about the attacker, the evidence channels, and the defender’s decision rule.

Concretely, we expect to:

  • implement a synthetic transcript generation environment with hidden harmful events or policy violations,
  • build baseline auditing policies such as fixed-threshold, random inspection, and simple evidence-aggregation policies,
  • run adversarial or semi-adaptive attacker strategies to see how policy performance changes,
  • compare policies under different noise levels, budgets, and signal strengths,
  • use the resulting experiments to identify failure modes and design principles for better auditing and escalation protocols.

If time permits, we would also like to explore a more realistic transcript-auditing benchmark or a lightweight interface for plugging in real transcripts later. However, the project is intentionally scoped so that the first useful version is a working simulation and evaluation framework, not a full production auditor.

Relevant work includes https://arxiv.org/abs/2409.07985, https://arxiv.org/abs/2606.06529, https://arxiv.org/abs/2512.07810, https://www.alignmentforum.org/posts/MDeGts4Aw9DktCkXw/toy-models-of-ai-control-for-concentrated-catastrophe, and work on constitutional classifiers. This project builds on the literature, but is more engineering-heavy.

Theory of change

Advanced AI systems are likely to be deployed with some mixture of monitoring, auditing, and escalation (see discussion surrounding AI auditors and verification measures). A major failure mode is that these mechanisms look good in theory but break under strategic behaviour, weak signals, or limited human attention. By building a testbed for auditing policies and running controlled experiments against adaptive attacks, this project can help identify which oversight strategies are robust, which are brittle, and where future research should focus. The resulting tools and findings should be useful for AI safety, AI security, and verification/escalation work more broadly.

Your role

Mentees will work as a small co-working research-engineering team split into different workstreams depending on the mentees' backgrounds. Each mentee should expect to own a concrete module or experiment slice, for example: environment design, policy implementation, attacker modelling, evaluation/visualisation, or writeup. We will work closely with the team, including pair-coding, design discussions, and regular feedback on implementation and experimental choices. The project will be structured so that mentees can make independent progress while still working toward a shared codebase and final report.

The project combines ML engineering, SWE, experimental design, and knowledge of AI safety. Different mentees may contribute by building the simulation environment, designing attacker and auditor policies, implementing evaluation infrastructure, analysing experimental results, or developing theoretical models of auditing. Applicants should have strong technical skills and be excited to learn about AI safety. Prior AI safety experience is helpful but not required.

Prerequisites

  • Strong Python skills and comfort working in a codebase independently
  • Ability to build and iterate quickly with the help of coding agents, e.g. Claude Code or Codex
  • Experience running experiments, debugging code, and interpreting results
  • Familiarity with basic probability/statistics and comfort reading technical papers
  • Basic AI safety research taste, good conceptual understanding of AI control and redteaming (which you can build by reading a bunch of posts and papers!)

Bonus: PyTorch, simulation environments, agent evaluation, or related SWE/ML engineering experience, formal study of game theory No deep game theory background required

Location preference

No local preferences. Ideally able to overlap with UK afternoon / early evening hours for weekly meetings and co-working.

Application question(s)

Please describe an experiment you would run to test whether an auditing policy is robust to an adaptive attacker. What would you implement first, and what result would count as informative? (300 words)

How much experience do you have with AI control? (e.g. how many papers have you read, talks you've attended, have you done any research projects in it?)

(Optional): Please share a coding or research sample, or briefly describe a project you built that involved experiments, simulation, model evaluation, or substantial software work. If relevant, include a link. (150 words)

About the mentors

Catherine Ge-Wang

Catherine Ge-Wang

University of Oxford

View profile

Cath is an AI safety researcher whose work focuses on AI control, scalable oversight, and mitigating emergent misalignment in increasingly capable agentic AI systems. They are currently a MIRI Fellow and ERA Fellow, where they are developing misalignment continuation evaluations for the UK AI Security Institute (UK AISI) and creating verification escalation protocols for frontier AI systems.

Cath is also a mathematics student at the University of Oxford, where they lead the Oxford AI Safety Initiative's Policy Team. Previously, they worked part-time at Redwood Research, and were a Fellow with MARS (mentored by Tyler Tracy at Redwood Research) and FIG (mentored by Liam Patell at GovAI).

L

Louis Thomson

Similar projects