Building a dynamic Bayesian protocol to rank which multi-agent risks the safety field should prioritize
About the project
The AI safety field is talent- and resource-constrained, so we need a principled way to decide which risks to defend against first. Multi-agent risks rank among the top concerns of AI experts (Saeri et al., 2026), yet there is still no systematic study of which inter-AI risks should be prioritized or which are most realistic. Given finite research capacity, which inter-machine threats warrant attention now, and how should that ranking change as evidence arrives?
Theory of change
Safety effort is currently allocated reactively. The field has no shared, evidence-weighted account of which inter-AI risks are most severe. As AI systems are increasingly deployed as interacting agents, this misallocation gets more expensive. Attention spent on speculative-but-vivid scenarios is attention not spent on the pathways most likely to actually materialize before transformative AI arrives.
Your role
I am seeking to create two sub-teams.
Team A would consist of experts in long-termism, and they would rank/construct specific scenarios of inter-AI collaboration. Ideally, I want them to be well-versed in Bayesian methods.
Team B would consist of empirical AI safety developers who would design toy models and scrutinize what kind of preconditions it would take for such risks to emerge in reality. I also expect them to compute timeline projections related to compute scaling. I expect the teams to work in a feedback loop and refine the protocol as more precise estimates become available.
Prerequisites
Both teams
- Bayesian methods
- AI control / x-risk familiarity
Team A
- Forecasting or elicitation experience
- Calibrated probabilistic reasoning
- Clear analytical writing
Team B
- Proficient in Python
- Trained/fine-tuned transformers (PyTorch)
- Familiar with scaling laws
Application question(s)
- What is the most realistic high risk scenario that involves multi-agents cooperation? (200 words)
About the mentor

I am most interested in how we can apply various mech interp devices to AI control. When it comes to inter-machine cooperation, I focus on what kind of monitor-evading communication is feasible (in particular, steganographic collusion).
Poseidon Research (current) - Monitoring for hidden reasoning SPAR '26 - Temporal Crosscoders Center for Human-Compatible AI '25 - SAEs to track social change
Northwestern University '26