Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

A Design Blueprint for Middle-Power AI Safety Institutes

International governance Technical governance National policy

Every state outside the US and UK is now told it needs an AI Safety Institute, but there is no design template scaled to a middle power's resources. This project produces a comparative anatomy of existing frontier-evaluation bodies and a modular, reusable blueprint a mid-sized state could adopt to build credible AI evaluation capacity without duplicating what larger institutes already do.

About the project

Independent capacity to evaluate frontier AI systems for dangerous capabilities is concentrated in a handful of well-resourced institutes (UK AISI, the US Center for AI Standards and Innovation, and a small number of others). Most states have neither the budget nor the technical staff to replicate them. There is no worked-out answer to the question a mid-sized government actually faces: what is the minimum viable institute that contributes something real to catastrophic-risk mitigation?

Core question: What should a middle-power AI Safety Institute actually do, cost, and look like — and how does it plug into the international evaluation ecosystem rather than reinventing it?

  1. Comparative anatomy. Mentees build a structured comparison of existing and emerging bodies - UK AISI, US CAISI, Canada's AISI, the EU AI Office, Japan's AISI, Singapore, and one or two others - coded along common dimensions: legal form and independence, mandate (evaluation vs. standards vs. market surveillance), budget, staffing and technical depth, relationship to government, and participation in international networks.

  2. Function mapping. Separate the functions that genuinely require sovereign capacity (evaluating nationally deployed systems, advising government on acute risk) from those better accessed through networks or shared infrastructure.

  3. The blueprint. Produce a modular design toolkit - an “AISI-in-a-box” - with budget tiers (a €10–15M lean model vs. a €50M+ model), staffing profiles, statutory-independence options, and a decision tree for what to build domestically vs. what to join.

Milestones: weeks 1–4, institution scan and analysis; weeks 5–8, completed comparative dataset and function map (midterm report); weeks 9–12, blueprint drafting and short paper (final report / Demo Day).

Theory of change

Credible, independent evaluation of frontier systems is a bottleneck for catastrophic-risk mitigation: dangerous capabilities can't be governed if few actors can measure them, and governance triggers (pause, restriction, disclosure) lack legitimacy without independent technical findings behind them. Today that capacity sits in a small number of countries, some of which have shifted rhetoric from “safety” toward “innovation” and “security.” Broadening the base of competent institutes increases the number of independent eyes on dangerous capabilities and strengthens the institutional substrate any future binding regime will depend on. A reusable blueprint lowers the cost for the next states considering this.

Your role

Mentees are the researchers. Each takes ownership of a subset of comparator institutions and, after we agree a shared coding rubric in the first two weeks, is responsible for populating it - reading founding legislation, budget documents, annual reports, and secondary analysis, and where feasible reaching out to people in the field. Beyond data collection, mentees contribute to the analytical work: the function-mapping and blueprint stages are collaborative, and I expect mentees to argue for their own conclusions about what belongs in a minimum-viable institute. Mentees will also co-author the final write-up. Autonomy is high on the descriptive work and more guided on the design synthesis, which we'll do together.

Prerequisites

Strong reading and synthesis skills across policy, legal, and institutional documents - this is a research-and-writing project, not a technical one.

A background in public policy, law, international relations, political science, economics, or a related field, OR demonstrated equivalent research experience.

Comfort working with primary sources (legislation, budgets, official reports) and turning them into structured comparative analysis.

Enough familiarity with the AI governance landscape to know what a frontier-model evaluation is and why it matters. Prior research experience welcome but not required. No ML or programming background required.

Location preference

No geographic requirement. A weekly meeting slot spanning European and North American time zones will be agreed with the team.

Application question(s)

(a) Pick any one existing AI Safety Institute or equivalent body. In your view, what is the single most important function it performs that a mid-sized country could NOT easily obtain by joining an international network instead - and why? Be specific. (max 300 words)

(b) A minister says: “We'll just give our existing market-surveillance regulator an AI safety mandate - why build a separate institute?” Give the strongest version of the minister's argument, then your best response to it. (max 300 words)

About the mentors

Michał Kubiak

Michał Kubiak

AI Safety Poland

View profile

Michał Kubiak is a researcher focusing on European AI regulations and AI risk management. His previous tech policy experience stems from his roles as an AI Policy Officer at the Observatorio de Riesgos Catastroficos Globales and at the Brussels-based European DIGITAL SME Alliance. Michał also has experience as a teacher at the AI governance bootcamps organised by ML4Good and at AI courses by BlueDot Impact, Electric Sheep, Sentient Futures and Tarbell Center for AI Journalism. Michał's earlier research experience includes industrial mathematics (STEM-based scientific problem solving for businesses, governments and other institutions).

Daniel Polak

Daniel Polak

AI Safety Poland

View profile

Daniel is a tax lawyer. His current research interests are internal AI deployment governance and the institutional design of national AI safety bodies. He also supports AI Safety Poland on policy projects.

Prior to this, he spent three years at KPMG working on M&A tax due diligence and international taxation.

Similar projects