Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Bayes-Optimal Research Protocols for Automated AI Safety Research

AI strategy Multi-agent systems Evaluations

We are designing optimal protocols for automated AI safety research. We use Bayesian models to determine how to best allocate compute across many AI agents doing open-ended research, so that safety research goes faster per unit of compute.

About the project

AI safety research is increasingly performed by AI agents themselves, and the binding constraint is shifting from researcher hours to compute allocation: how many agents to run, on which problems, in parallel or sequence, replicating or exploring. This is a resource-allocation problem under uncertainty, and it has a normative theory. Optimal experimental design, bandit and Bayesian search models, and formal metascience (models of optimal science funding and the explore-exploit structure of collective inquiry) tell us how a Bayesian planner should allocate resources across an open-ended research portfolio.

This project develops that theory for the automated-research setting. Questions include: (i) optimal allocation between exploring new hypotheses and verifying or replicating existing findings, when verification is cheap but correlated errors are possible; (ii) how allocation rules should respond to the distinctive error structure of AI researchers (shared blind spots across instances of the same model, sycophantic convergence); (iii) when parallel independent runs beat sequential ones with shared memory; (iv) stopping rules — how much compute a research direction should get before being abandoned. The theory of change: labs are already making these allocation decisions ad hoc; principled protocols could measurably accelerate safety research per unit of compute.

We will test the theory empirically in a sandbox: a population of models doing automated theorem proving, a domain with verifiable ground truth and tunable difficulty. Running competing allocation strategies on the same problem sets lets us measure which strategies perform best and the magnitude of gains from strategic compute allocation via our solutions.

Deliverable: a formal model with analytic results where tractable, the theorem-proving sandbox with head-to-head strategy comparisons, and a paper aimed at alignment workshops or a formal epistemology venue.

Theory of change

If successful, our protocols will tell labs how to allocate compute across AI agents doing safety research to maximize research output, with empirically measured gains from our theorem-proving sandbox — accelerating every safety agenda that depends on automated research.

Your role

Mentee autonomy is a core value for me: I set the research direction and modeling framework, but I expect mentees to shape the project, and their input and innovations are genuinely welcome — the best version of this project includes ideas I haven't had. Mentees will own concrete workstreams: formalizing allocation problems, proving results in simple regimes, building the theorem-proving sandbox, and running head-to-head comparisons of allocation strategies in it. Everyone drafts sections of the write-up with detailed editorial feedback from me. Expected trajectory: tightly scoped tasks in weeks 1–3, increasing independence thereafter.

Prerequisites

Required: comfort with probability and expected-utility reasoning; working Python. Helpful but not required: exposure to bandit problems or optimal experimental design (I will teach this); experience running LLMs via API or agentic scaffolds; familiarity with automated theorem proving or proof assistants (e.g. Lean) for the sandbox track.

Location preference

No geographical requirement. Mentees must be available for a weekly team meeting on a weekday between 13:00 and 22:00 UTC (9am–6pm US Eastern). This window works well for mentees in the Americas, Europe, and Africa; mentees in Asia and Oceania are welcome if a late-evening or early-morning slot work

Application question(s)

  1. You have 100 units of compute and 10 candidate research directions with unknown success probabilities. Sketch a Bayesian allocation policy and explain how it differs from splitting compute equally. What changes if two directions' outcomes are highly correlated? (300 words)

  2. Name one way a fleet of AI agents doing research differs from a population of human scientists in a way that should change optimal allocation, and how. (150 words)

  3. Link to each a writing sample and code sample.

About the mentor

Aydin Mohseni

Aydin Mohseni

Carnegie Mellon University

View profile

I am an assistant professor of philosophy at Carnegie Mellon University, a core member of CMU's Institute for Complex Social Dynamics (ICSD), and a founding member of CMU's Conceptual Foundations of Safe AI initiative (CoSafe). My research applies game and decision theory, Bayesian statistics, and evolutionary models to questions in metascience and AI safety. Recent work includes a decision-theoretic account of when it is rational for an agent to pause and refine its values before acting (with Alex John London), a Bayesian reduction of causation in causal models (with Daniel Herrmann, Benjamin Levinstein, and Bruce Rushing), and work on AI alignment as a principal-agent problem. In 2026 I am organizing a workshop at CMU on the foundations of AI agency & interpretability.

Similar projects