In this project, we'll do deconfusion research on acausal interactions and related strategic dynamics in order to enable better outcomes from interactions between ASIs.
About the project
The scope of this project is pretty broad, and the specific topics we focus on depend on mentee backgrounds and interest. While the primary focus of the project is acausal interactions, some work might be relevant to causal interactions as well.
Here are some sample possibilities:
Acausal trade: Most discussions of acausal trade presuppose two agents in fairly similar worlds trading with each other. But there are many more possibilities for what the acausal trade landscape might look like, including indirect trades brokered by trusted third-parties, trades between worlds with large asymmetries in simulation capability, etc. Exploring underinvestigated possibilities could be important to inform new interventions and identify blind spots in our thinking about these topics.
Stability of acausal coalitions / cooperation schemes.
Decision-making: How agents make decisions at a low-level could affect how they approach some acausal interactions. In particular, I've recently gotten pretty interested in how architectural choices in how an agent's identity (ie source code) is exposed to it could affect its strategic choices.
Commitment races: Commitment races are the main mechanism why powerful AIs might have catastrophic conflicts with one another, despite this being ex post disvaluable to all parties. Formalising high-level models of commitment race dynamics and exploring how much of an obstacle CRs are to cooperation strategies like Safe Pareto Improvements could be pretty important.
AI Metacognition: How should AIs think about these topics? Are there pitfalls we could steer AIs away from?
Theory of change
The main goal of this project is to enable better outcomes from interactions between ASIs by informing interventions on how AIs are disposed to handle activities like acausal trades. Such interventions might look like making certain information salient to an AI system early in its training or deployment, or recommending low-cost architectural changes.
Note: This isn’t a theory of victory per se. Ensuring that all AI systems deployed on earth act optimally in all interactions involving other powerful agents is unlikely to be sufficient to prevent extinction and to bring about a positive future.
However, it’s possible that much more progress towards improving AI-AI interactions is necessary to prevent bad futures. For example, an aligned ASI could wind up expending all its resources in a zero-sum conflict with an unaligned or alien ASI.
And even if we fail to solve alignment and a good outcome within our lightcone is not on the table, it might still be impartially very valuable to work on this in order to help aliens who might share some of our values benefit from interactions with earth-originating AIs.
Your role
Mentees with a strong formal background could help formalise different models of things like decision-making under uncertainty about one's logical identity, coalitions, strategic proxies. Mentees can propose and investigate new approaches to understanding acausal interactions and related strategic dynamics. Mentees can critique previous work in this area.
The plan is to give mentees a good basic understanding of this research area, and then help them figure out what they're most interested in exploring in greater depth. If there's more than one mentee we will have weekly group discussion meetings in addition to 1-1s with me.
Prerequisites
A decent game theory background - taking this course should be sufficient (https://www.coursera.org/learn/game-theory-1), though more may be helpful
Prior familiarity with updateless decision theory. Can be gleaned by reading a bunch of blog posts on LessWrong (https://www.lesswrong.com/tag/updateless-decision-theory)
Application question(s)
About the mentor

James researches acausal interactions and AI conflict at Center on Long-Term Risk. Recent topics he's been working on include understanding commitment races in terms of strategic proxies and the effect on logical correlations of how agents identify what algorithm they are running. He's also been working on a better UI for interacting with AI agents doing autonomous work.