Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Convergent Truth

Multi-agent systems Behavioral evaluation of LLMs Scalable oversight

Design a test for whether LLM agents, given opposing information and sometimes goals, can still converge on the truth.

About the project

We increasingly rely on AI agents: to compile research, write code, and provide perspectives. But as those agents have improved at persuading us, it raises the question of whether they are doing so well, especially whether they are resistant to misleading, false, and deceptive evidence. That is, how can we characterize whether agents engage in good versus bad persuasion?

To do so, we put agents in a collaborative environment where they must rely on and provide evidence to each to come to the correct answer. Each agent may have different information and a different reward function from the others; a few succeed by sabotaging the group but most want to arrive at the truth. Thus these agents must, like the blind men in the parable of the elephant, work together to find the truth given their partial and conflicting knowledge.

We first evaluate a toy domain where agents see the outcomes of pairwise matches of tug-of-war and, from these, must infer the strength of the (separate) competitors with the complication that competitors can be lazy and throw a match.

We then extend this to a naturalistic setting in which agents with different access to information about a pull request must work together to match the human ground truth issue.

Firstly, we study the behavioral strategies which emerge through interaction, such as when they reveal information, lie, incorporate new evidence, and come to agreement. We also measure whether agents arrive at the rational, Bayesian solution.

Secondly, we train agents to be responsive to their reward functions, studying what kinds of persuasion emerge and whether agents so trained generalize to downstream evaluations.

Our (possible) findings

  • How different groups agents behave in semi cooperative persuasive environments
  • Whether they are robust to sabotage and
  • How well agents approximate the rational solution

Theory of change

We are at a path-dependent juncture in AI. The research directions we take now, if chosen wisely, will help develop technical and conceptual countermeasures against AI takeover and other risks; they will reduce the "alignment tax." While there has been an increasing amount of attention paid to AI safety, proportionally few are working technically on value alignment despite the fact that to align AI systems we may need to make them understand values and morality like humans do. I believe that work in value alignment, both because it is a relatively under-explored direction with room for more formalization and because it clearly relates to making AI do what we want, will most effectively reduce the alignment tax. In particular, robust formal accounts of persuasion are necessary to diagnose and mitigate models’ deceptive capabilities. Having early warning signs of deceptive capabilities may actually help focus the attention of society in a way to best mitigate issues.

Your role

Depending on time availability and level of experience mentees will either lead or co-lead the project. This will involve:

  • Designing/tinkering with the persuasion game protocol.
  • Establishing (and formalizing?) a rational Bayesian bound on the persuasion game
  • Running and iterating on experiments to make sure LLMs can play the game well and that it appears to be measuring what we want.
  • Scaffolding real world data to have similar game dynamics; testing the generalizability of the game dynamics in other settings.
  • Training agents using the game success as a reward.
  • Co-writing the paper.

Prerequisites

Prior research experience (can be in another domain).

Able to read and understand papers in at least some AI domains (and not just use an LLM summary)

Proficiency with python. Writes clean, understandable code.

  • It is fine to use coding agents but you must know how to reign them in (viz.: read their code; don't let them reimplement everything; use linters and formatters).

Have a realistic model of your own schedule and get things done when you say. (I'm not trying to say everything has to be done quickly -- just that you should communicate your uncertainty.)

Openness to learn new things.

Clear communicator. Agreeable and humble

Application question(s)

In a few sentences each...

  1. What does a good research collaboration look like to you?

  2. What's a book or piece of media that changed or inspired you? Why?

  3. Tell me about a project you have worked on. (Ideally a research project.) What role did you play? What went well? What didn't?

About the mentor

Jared Moore

Jared Moore

Stanford University

View profile

Similar projects