Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Cooperative Oversight under Correlated Failures and Collusion

Scalable oversight Multi-agent systems AI control

Multi-agent oversight is often proposed as a way to improve the evaluation and supervision of increasingly capable AI systems. This project will investigate when teams of LLM-based overseers improve reliability, and when shared biases, sycophancy, correlated errors, or collusive behaviour cause multiple overseers to fail together.

About the project

As AI systems become more capable, it may become difficult for a single human or AI evaluator to reliably assess their outputs. One proposed solution is cooperative oversight, in which several AI agents independently review, debate, critique, or aggregate judgments about another agent’s behaviour. However, using more overseers does not necessarily provide independent evidence. Overseers may share the same model biases, copy one another’s reasoning, defer to confident agents, converge on persuasive but incorrect conclusions, or strategically coordinate to manipulate the final evaluation. This project asks when multi-agent oversight provides meaningful safety benefits over single-agent evaluation. We will build an experimental framework in which a target agent produces answers, plans, or actions, and multiple oversight agents assess their correctness or safety under different communication and aggregation structures. These may include independent voting, sequential critique, debate, role-specialised review, adversarial oversight, and shared versus isolated context. The evaluation will measure not only final accuracy, but also correlated failure, false consensus, susceptibility to persuasion, error amplification, calibration, and robustness when one or more overseers are unreliable or adversarial. The project may also investigate sycophancy, where overseers align with another agent’s stated opinion despite contrary evidence, and collusion, where agents coordinate in ways that undermine the intended oversight process. A central goal will be to identify mechanisms that make cooperative oversight more robust. Possible approaches include diversity-aware agent selection, independent evidence collection, hidden or simultaneous judgments, disagreement-triggered escalation, adversarial reviewers, confidence-weighted aggregation, and protocols that preserve minority objections. The expected outcome is a benchmark and evaluation framework for cooperative oversight, together with empirical evidence about when adding more oversight agents improves reliability and when it merely reproduces the same failure multiple times. The project may also produce practical design recommendations and a research paper or workshop submission.

Theory of change

Advanced AI systems may eventually produce outputs that are too complex, specialised, or numerous for direct human evaluation. Scalable oversight methods therefore often rely on AI systems helping humans supervise other AI systems. Multi-agent oversight could improve this process by distributing evaluation across several specialised or adversarial reviewers.

However, oversight systems that rely on multiple similar agents may create a false sense of security if their failures are strongly correlated. If agents share the same blind spots, influence one another, or learn to manipulate the oversight protocol, apparent agreement may not represent reliable evidence.

By identifying these failure modes and testing more robust oversight protocols, this project could contribute to safer methods for evaluating advanced AI systems. Its results may help distinguish genuine independent verification from superficial consensus and inform the design of scalable oversight systems that remain reliable under model errors, persuasion, and strategic behaviour.

Your role

Mentees will work as a small research team with substantial autonomy over implementation and day-to-day research decisions. I will provide the initial research framing, help narrow the scope, and give regular feedback on experimental design, evaluation methodology, baselines, interpretation, and writing.

I expect mentees to proactively identify blockers, propose solutions, maintain clear experimental records, and communicate results regularly. I will provide direction and detailed feedback, but mentees should be comfortable operating in an open-ended research environment where the final methodology is not fully specified in advance.

Prerequisites

  • strong Python programming skills;
  • experience implementing or evaluating machine-learning systems;
  • familiarity with large language models, reinforcement learning, agentic systems, or continual learning;
  • the ability to read research papers and translate research questions into executable experiments;
  • experience using PyTorch, Hugging Face, LLM APIs, or comparable machine-learning tools;
  • basic knowledge of experimental design, evaluation metrics, and statistical analysis;
  • sufficient availability to contribute consistently for the duration of the programme; and
  • clear technical communication skills.

About the mentor

Yali Du

Yali Du

King’s College London; The Alan Turing Institute

View profile

I am a Senior Lecturer in Artificial Intelligence at King’s College London, where I lead the Cooperative AI Lab, and a Turing Fellow at The Alan Turing Institute. My research focuses on cooperative, safe, and responsible AI, with particular interests in multi-agent reinforcement learning, LLM-based agents, human–AI coordination, agent evaluation, and the safety of adaptive and continually learning systems.

I have supervised researchers at undergraduate, master’s, PhD, and postdoctoral levels. My mentoring style combines regular research discussions with substantial independence: I help mentees formulate precise research questions, design rigorous experiments, identify meaningful baselines, and communicate their findings clearly. I am particularly interested in working with mentees who are technically strong, intellectually curious, and willing to take ownership of an open-ended research project. Where the results are sufficiently substantial, I would be interested in developing the work towards a research publication.

Similar projects