Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Evaluating Debate Protocols in Auditing Sabotage Bench

Scalable oversight AI control

"Auditing Sabotage Bench" is a dataset of slightly sabotaged papers and codebases, which (after malicious editing) have very different results than the original paper. Can we use classic debate protocols to make humans (and LLM judges) better at identifying research sabotage, and uplift human auditors?

About the project

In "Auditing Sabotage Bench" (https://arxiv.org/pdf/2604.16286), Gan et al showed that humans are quite bad at identifying research sabotage! That is, when shown either (1) a real research codebase, or (2) a subtly sabotaged codebase, with malicious changes to change the results of the experiments, humans have a hard time differentiating the two. This seems like a perfect testbed for scalable oversight (https://www.lesswrong.com/w/scalable-oversight) — can we use debate to improve the abilities of humans or LLMs to identify sabotaged research?

This will largely be the application of previous empirical debate research (e.g. https://www.lesswrong.com/posts/x59FhzuM9yuvZHAHW/scaling-laws-for-scalable-oversight, https://arxiv.org/pdf/2407.04622) to this new setting. However, Auditing Sabotage Bench would arguably be the most complex setting yet for scalable oversight, and I'm very interested about our potential to observe if the dynamics of debate change as we move from toy examples to something much closer to the actual threat model (superhuman AIs doing inexplicable research which we can't verify).

There is a lot we can do here, but the central question is this: can we use debate protocols to improve judge performance on Auditing Sabotage Bench?

Some directions we can take this:

  • Which debate protocols perform best here?
  • How do the dynamics change if we have a smarter model argue that some sabotaged research is actually correct, and a dumber model argue that it might be sabotaged? (mimicking the case in AI control (https://www.alignmentforum.org/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion), where our smartest model might be misaligned, but we have an array of less-smart aligned models)
  • If we have a fixed judge/debater budget, how should we optimally spend that budget to improve the odds of the judge getting the correct answer?

Theory of change

As AIs get smarter, and AI safety researchers rely more on AI, we run into the issue that our AIs might be subtly sabotaging our research (https://blog.redwoodresearch.org/p/how-can-we-solve-diffuse-threats). This gets quite bad if our alignment plan depends on research done by superhuman AIs, or relies on enough research from AIs that it becomes difficult to verify the correctness of all of it.

The classical answer to problems of supervising an AI that's smarter than us is scalable oversight—in particular, we could try having our superhuman AIs debate, and hope the dumber human judges learn true things via this process. Previous research has had mixed results (https://arxiv.org/abs/2407.04622), but the actual usefulness in really tricky situations is an open question.

With the release of "Auditing Sabotage Bench" (https://arxiv.org/abs/2604.16286) I think we have a much better testbed than ever before to see when and whether debate can be useful. I hope that experiments along these lines can inform how we choose to evaluate research done by untrusted AIs, which I anticipate to be pretty crucial during the critical period.

I'm quite exciting about lots of things in this category, and have previously worked on human auditing for AI Control (https://www.lesswrong.com/posts/2CJyfgaJQk8pyRSCp/auditing-1)

Your role

Mentees will have significant amounts of autonomy. I will not be actively working on this project outside of this mentorship (although this mentorship would include 1-2 meetings a week, and a slack channel for more frequent comms). See the project description for the first steps / questions I have here, but I expect the exact thrust of the project to change based on your interests, and how early experiments come out. (Also, I'd encourage you to e.g. try out variations of the experiment if you think we might find something interesting, although I will probably be opinionated about what sorts of experiments I'll find a priori promising.)

Mentees will be the primary researchers on this project. I expect most of my involvement to come in the form of frequent feedback about roadblocks / experiments / analysis, and also periodic discussions about new directions and what seems promising.

Prerequisites

  • Proficient programmer, or at least can verify AI-assisted code
  • Can read and understand "Auditing Sabotage Bench" (https://arxiv.org/pdf/2604.16286)
  • (Highly preferable but not strictly necessary) Experience with ML research

Location preference

Preference for UTC-8 to UTC+2, although we could likely figure something out if not.

Application question(s)

  1. Read or skim Kenton et al (2024), https://arxiv.org/pdf/2407.04622 . With special attention to the example debate on the last page, please explain one or more ways this debate would be different in our case — if the debaters were trying to convince a judge that some particular research codebase was or was not slightly sabotaged by bad design or implementation choices. (<300 words)

  2. Provide links to one or more past research projects that you're proud of. This can be both an empirical or a conceptual project, and it doesn't have to be an AI safety project. If all of your best projects are private, provide a short description of the project and your role in it. (<100 words)

  3. (Optional) List some questions you have after reading the research proposal, or some experiments you might be excited to run in this direction. (<300 words)

About the mentor

Joey Yudelson

Joey Yudelson

Aether Research

View profile

Joey is a technical AI safety researcher working at Aether Research, currently working on training dynamics. He was previously a MATS 7.0 scholar with Ryan Greenblatt at Redwood, working on reward hacking and character misgeneralization.

Similar projects