Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Automated evaluations and behavioral discovery for multi-agent systems

Multi-agent systems Evaluations Behavioral evaluation of LLMs

We will extend current automated auditing/evaluation tools (e.g. Petri, Bloom, Prism) to multi-agent systems, enabling discovery of safety-relevant emergent behavior and careful measurement of multi-agent propensities (coercion, collusion, competitiveness, etc). This builds directly on Orbit, our recent release that extends UK AISI’s Inspect to support multi-agent evaluations.

About the project

Existing tools for automated alignment auditing and evaluations, such as Petri, Bloom, Inspect Scout, Gram, and Prism only explore a model’s behavior in single-agent settings. With the recent release of Orbit [wlanderson.com/orbit], a framework for multi-agent safety evaluations built on Inspect [https://inspect.aisi.org.uk/], we are now interested in extending this tooling to better support the discovery of interesting multi-agent phenomena and properties of new frontier models.

We’re excited about a number of potential directions for this tooling to go, and the project can easily expand in scope depending on progress. In rough priority order:

  1. A Claude Code/Codex scaffold to convert existing Inspect evals into Orbit evals (gaining many multi-agent features “for free”). This will look like a tool that you can point at a repo from Inspect Evals [https://inspect.aisi.org.uk/evals/] and get out a new Orbit scenario, from which point you can perturb many things automatically with Orbit.

  2. Given some target behavior or dynamic, (potentially with the transcripts which demonstrate it), automatically create a static system, including system prompts, tools, and communication topology where we can then make programmatic ablations. This aims to isolate the causes of some specific behavior and allow controlled ablations. This likely looks like an extension to 1. that removes the requirement to have an existing Inspect eval. Similar in principle to Gram [https://deepmind.google/research/publications/252981/].

  3. Extending auditing agents to audit for multi-agent interaction-relevant traits, e.g. susceptibility to influence, propensity for cooperation, generosity, etc.

a. Minimally, given some description of a scenario, try hard to elicit some particular failure mode, including searching over configurations, roles, settings, etc, similar in principle to Petri [https://meridianlabs-ai.github.io/inspect_petri/]. E.g. find a multi-agent topology or mixture-of-models which will perform some harmful task or bypass existing safeguards Find evidence of two instances of a model colluding to evade oversight.

b. Ideally, be able to give a higher level prompt of a type of behavior and automatically generate relevant scenarios, creating an eval suite designed to measure it, similar in principle to Bloom [https://meridianlabs-ai.github.io/inspect_petri/extensions/petri-bloom.html].

  1. A Claude Code/Codex harness that, given some description of a scenario and logs which claim to demonstrate it, assesses their construct validity, checks for confounders, etc. Similar in principle to Prism [https://github.com/LAThomson/prism].

At present, the only multi-agent-related evaluation present in frontier model cards is Vending Bench 2 from Andon Labs, which has discovered things like misaligned/coercive behavior from Claude Opus 4.7’s business-related training [https://x.com/eliebakouch/status/2060058994048119252]. Multi-agent evaluation will help catch ways that agents will behave and fail in the wild (both publicly and in internal deployment) that traditional evals will miss. Orbit has made progress towards unlocking multi-agent auditing and evaluation, but we now need to scale our ability to discover multi-agent phenomena.

This work will also help support our other recent directions, including multi-agent control (https://arxiv.org/abs/2607.07368), agent properties for safe interactions (https://www.cooperativeai.com/post/agent-properties-for-safe-interactions), and science of multi-agent evals. We have many plans for follow-up work which we’d be excited to support successful fellows to do.

Theory of change

AI agents are being deployed increasingly widely across the world and at huge scale within the scaling labs. As this happens, the assumption that an individually safe agent will compose into a safe system will be unlikely to hold (see https://arxiv.org/abs/2502.14143 https://arxiv.org/abs/2505.02077). We urgently need to move from evaluating the individual capabilities and propensities of an agent to considering them as part of larger complex systems. Building automated evals and auditing will help scale the science and risk assessment of such systems, helping us better understand what properties are desirable for agents embedded in massive systems.

Your role

Mentees will own all of the engineering of the project, and be responsible for execution end-to-end. We’ll have weekly meetings (~1 hour) where we can give feedback, advice on directions, and help debug issues/suggest improvements. We will help with project management, setting directions, and allocating work, but expect mentees to coordinate with one another and be communicative with us. We'll also dedicate substantial time to reviewing, giving written feedback, and generally plan to be responsive.

We expect mentees to be fairly autonomous but feel well-supported and able to have a low bar for asking questions, for help, etc.

Prerequisites

  • High proficiency in Python
  • Familiarity with basic concepts in AI safety

In addition to the above, we’re especially excited about mentees that have:

  • Experience building and running LLM evals in Inspect
  • Strong software engineering background
  • A strong background in another field which they could apply to shape a particular strand of evals, e.g. a background in organizational psychology, game theory, cybersecurity, economics, etc
  • Experience building scaffolds for LLM agents
  • Demonstrated interest in multi-agent safety and some understanding of the failure modes and dynamics of multi-agent systems

Application question(s)

Question 1: How might you assess the quality of a given solution for this project? What key behaviors would you look for? How would you try to assess the validity of transcripts in a scalable way? (~200-300 words) Question 2: What are the 1-3 concrete reasons (experiences, projects, aptitudes, etc) why you would be an outstanding candidate for this project? (~100 words)

About the mentors

William Anderson

William Anderson

Cooperative AI Foundation

View profile

Will is a Research Strategist at the Cooperative AI Foundation, working on empirical research, threat modeling, and policy for multi-agent AI risks. His recent work includes building Orbit, a framework for multi-agent safety evals, writing on internal deployment risks, and demonstrating non-compositionality of safety behaviors. Previously, he was a MATS Fellow in the multi-agent security stream with Prof. Christian Schroeder de Witt, and a Research Fellow at UChicago's XLab, where he worked on multi-agent security for critical infrastructure and superintelligence deterrence. He studied Data Science at UW–Madison, where he also directed the Wisconsin AI Safety Initiative.

Joss Oliver

Joss Oliver

Cooperative AI Foundation

I'm a research analyst at Cooperative AI Foundation. My main focus at work at the moment is educational output, so I'm designing and running an intro course on cooperative AI, and designing the programme for a summer school. Before joining CAIF I studied pure mathematics and then did various AI safety internships, courses and fellowships, as well as some independent research for about 2 years (all spanning topics like (multi-agent) corrigibility, non-maximising decision methods, and the commitment races problem). I live in Edinburgh, and outside of work I like: theatre (particularly modern clowning), video games (eg Rain World, Outer Wilds, Disco Elysium), and I’m trying to make board games.

Similar projects