We investigate through simulations how effective different AI policies would be in reducing AI risk. We will collect a dataset of existing and proposed AI legislation, and create an agentic simulation of the world (countries, companies, etc), iteratively increasing in complexity throughout the project.
About the project
AI governance policies will likely be necessary to allow companies and countries to coordinate on developing AI safety, e.g. by not rushing to release models before adequate testing, or holding to mutually agreed-upon compute limits [1] for new pretraining or RL runs. However, the people writing such policies may not be very familiar with the technicalities of AI alignment. We aim to create a system to simulate the effects of proposed AI policies on the world, to enable policy researchers to choose policies more likely to be effective at increasing safe deployments, or avoiding catastrophic outcomes from our default path of no interventions.
Our project consists of two parts. First, existing and proposed AI policies will be collected along with various bills that passed (or were proposed) in the US and other countries; starting points include OECD.AI and California's frontier-model legislation. This will give an idea of what type of interventions we might consider, and we will turn it into a useful dataset for simulation. Second, an agentic simulation will be created, initially based on the actors in the AI 2040 [2] scenario (companies, governments, etc). We will start with a simple simulation with basic economic and technological progress variables, and iteratively increase the complexity of this simulation throughout the project.
This simulation is meant to surface different possible outcomes that may not have occurred to the user. The aim is on exploratory breadth rather than prediction accuracy, which is nearly impossible to achieve on such a large scale. Overall, our simulation framework helps researchers enumerate possible future outcomes of AI, and enables them to investigate their own specific policy questions. It will also accelerate the creation of public wargames [3] in the future to educate the public.
[1] https://arxiv.org/abs/2402.08797 [2] https://ai-2040.com based on https://ai-2027.com [3] https://www.intelligencerising.org/
Theory of change
Our project helps identify which policies would have an actual impact on risk, and whether or not current AI policies are on track to have this effect.
Your role
Mentees will have substantial autonomy and there are several possible types of work, from dataset building to policy identification to agentic coding. See proposal.
Prerequisites
- Highly proficient with Python and git (or Github).
- Has written code to run transformers from Huggingface, with custom system prompts.
- Ideally, a background in AI governance or legal studies, or an interest in geopolitics.
Application question(s)
In your opinion, how long will it be until nearly all coding tasks can be automated by AI? Why do you say this timeline? Optional reading: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ (200 words)
What is one way in which the scenarios in AI 2027 (or in AI 2040) are unrealistic? What aspect do you think is most likely to come true? https://ai-2027.com and https://ai-2040.com (300 words)
About the mentors

David lives in Canada (and often the UK) and enjoys AI safety research and mentorship. He works full-time as a Senior Research Manager at ERA, helping fellows execute research projects in each ERA AI fellowship. He also co-runs Lida Safety Research. Previously, David participated in the MARS program with Geodesic Research, conducted research at Mila, and was an early member of Yoshua Bengio's LawZero in Montreal. He was lead instructor at AI Security Bootcamp Singapore, and will serve the same role at the upcoming event FAST (Frontier AI Security Training) in Singapore.
David has previously mentored for SPAR, Athena, and other mentorship programs. He spent four years as a cyber insurance startup CTO, leading a 15+ person team. He holds a cybersecurity PhD from Columbia University and worked with 25+ students overall. David also works in AI risk communications, with a 30,000+ subscriber YouTube channel.

Linh is an independent AI safety researcher at Lida Safety Research, which she co-runs. Her research focuses on simulating AI governance policies to determine their effectiveness and feasibility. Previously, Linh worked on alignment at Mila through latent adversarial training for personalities. She participated in the MARS research program with Geodesic Research, working on chain of thought monitorability. Earlier, she did a postdoc at the University of Technology Sydney and obtained her PhD from the University of Queensland in natural language processing. Linh was a SPAR mentor in a previous iteration, and enjoys hackathons: she has won 1st place twice and 4th place once at Apart Research hackathons. She also won 4th place at the Redwood Research Alignment Faking hackathon.