Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

The Commitment Atlas: mapping AI red lines and testing whether they can be verified

International governance AI strategy Technical governance

Governments, labs, and scientists keep declaring AI red lines, but no one has mapped them or tested whether they can be checked.

Mentees will build a public, sourced dataset of these commitments and convert the strongest into draft verifiable standards.

About the project

Red lines are multiplying. IDAIS Beijing proposed four in 2024: no autonomous replication, no power-seeking, no weapons-of-mass-destruction uplift, and no autonomous cyberattacks. Twenty companies signed the Seoul Frontier AI Safety Commitments. The Global Call for AI Red Lines gathered over 300 signatories, including 11 former heads of state (https://red-lines.ai). It launched at the UN General Assembly in September 2025 and demands enforceable red lines by the end of 2026. The Athens Roundtable identified seven red line domains in December 2025.

Two things are missing:

  1. No dataset spans these commitments: METR, SaferAI, and the Future of Life Institute track company frameworks. Nobody tracks state, corporate, and international commitments together, coded for whether they could be checked.
  2. Most commitments cannot be checked as written: On bioweapons uplift, Anthropic sets its threshold at basic STEM knowledge, OpenAI at novice actors, Amazon at expert level. Companies keep about half of their voluntary safety pledges. Most safety assessments are still done by the companies being assessed (UN scientific panel, July 2026).

This project builds the Commitment Atlas. Every public AI red line commitment from a government, company, or international process goes into one structured dataset. Each entry is coded for actor, instrument, domain, trigger, threshold precision, verification hook, and consequence. Every cell traces to a primary document.

The Atlas answers three questions nobody can answer today. Where do all major labs, and both the US and China, already agree? Which commitments are precise enough to test? Which carry any consequence at all?

Then we stress test the hardest step. Two case studies each take one commitment from the overlap set. Each drafts what a testable standard would look like: definitions, evidence, who could check. A short precedent review asks what FATF, WADA, INPO, and maritime load lines teach. That includes where FATF produced paper compliance instead of results.

Everything publishes openly with named mentee credit. The findings feed the founding design of the Touchstone Council, an independent verification institution I am building. The round ends weeks before the Global Call's own deadline. That timing is the point.

The attached proposal has the draft codebook, the week-by-week plan, and full references. It builds on Thwing et al. 2026 (https://papers.ssrn.com/sol3/papers.cfm?abstract\_id=6854278) and departs from it on institutional form.

Theory of change

Verification is the binding constraint on red lines. IDAIS Venice named it in 2024. The Athens Roundtable reached the same conclusion in 2025: turning norms into checkable thresholds is the hard part. If commitments cannot be checked, they will not hold. The Global Call's end-of-2026 deadline will then pass with nothing behind it.

This project supplies two missing things. A sourced map of what has actually been promised. A tested method for converting one promise into a checkable standard. Thwing et al. 2026 documents why the intergovernmental measurement network cannot yet carry verification. We share that diagnosis and depart on the prescription. The Touchstone Council, which this work feeds, is designed as an independent body funded by no one it assesses.

The theory of change is concrete. The Atlas identifies the overlap set of red lines every major lab and both the US and China already accept. That set is where verification work can start without a treaty. The case studies show whether conversion to testable standards works at all. Both outputs go to the Council's advisers and into the governance community I work in.

Your role

Each mentee owns one workstream end-to-end.

Seat 1: state and international commitments. Collect and code every government and multilateral red-line instrument.

Seat 2: company commitments. Collect and code the frontier safety frameworks and related pledges. Familiarity with model evaluations helps here.

Seat 3: precedents and standards. Lead the FATF, WADA, INPO, and load line review, and co-draft the two case studies.

With a fourth mentee, the case studies become their own seat.

Autonomy is real. You make the coding calls in your workstream and defend them in our weekly adjudication. You draft your sections of the published outputs. I review everything, weekly, in writing. Hard cases we decide together. You are a co-author, not a research assistant.

Prerequisites

  • Strong analytical writing in English. You will be judged on a sample.
  • Research experience in policy, law, international relations, or a nearby field. Serious coursework counts.
  • Careful sourcing habits. Every cell in the Atlas must trace to a primary document.
  • Comfortable structuring data in spreadsheets. No coding and no machine learning background required.
  • Eight reliable hours per week, mid-September to mid-December.

Location preference

None

Application question(s)

  1. Pick one public AI red line commitment from any government, company, or declaration. Link it. What observable evidence would establish a violation, and who could observe that evidence today? (250 words)

  2. Thwing et al. 2026 argue that the international network of AI safety institutes should grow into the verification role: https://papers.ssrn.com/sol3/papers.cfm?abstract\_id=6854278. Read sections I and III only. What is the strongest objection to that claim? (200 words)

  3. Link one writing sample, ideally analytical or research writing.

About the mentor

Aryan Agarwal

Aryan Agarwal

Touchstone Council (founder); OECD (Policy Analyst, applying in a personal capacity)

View profile

I am a Policy Analyst at the OECD Paris HQ, where I work with the Science, Technology and Innovation Directorate. In my role, I have led and negotiated texts for 3 ministerial meetings and worked on diplomatic outreach. Before policy, I co-founded a regulated fintech.

I am now building the Touchstone Council, an independent institution that turns declared AI red lines into testable standards and verifies who keeps them. Its design borrows from FATF mutual evaluations, WADA lab accreditation, and IFRS standard setting.

As a mentor, I run a weekly team call, a short one-to-one with each mentee, and give written feedback on every draft. Mentees get named credit on everything we publish, plus access to my network. I am not a ML researcher, and this project does not require you to be one. It needs careful readers who write plainly and source every claim.

Similar projects