Compute governance proposals assume detection probabilities that nobody has estimated. This project builds the first sourced evidence base on what inspection instruments actually detect, using the nuclear safeguards record as the comparison case, then uses those numbers to say which compute enforcement architectures can work and which cannot.
About the project
Compute is the most tractable object of frontier AI regulation because it is countable, physical and concentrated (Sastry et al., 2024, arxiv.org/abs/2402.08797). Nearly every proposal built on that observation, from training-compute thresholds to licensing and tradable permits, requires a regulator to observe how much compute a developer used and to detect under-reporting with some probability. The technical literature describes what could be built: on-chip governance mechanisms (Aarne, Fist and Withers, CNAS, January 2024, cnas.org/publications/reports/secure-governable-chips), location verification (Brass and Aarne, IAPS, 2024, iaps.ai/research/location-verification-for-ai-chips), training-run monitoring (Shavit, 2023, arxiv.org/abs/2303.11341) and hardware-enabled mechanisms for export control (Kulp et al., RAND WRA3056-1, January 2024). What the field does not have is an evidence base. How often would each instrument catch a determined evader? What does a false alarm cost, and who bears it? How does a regime perform once the regulated party optimises against it? In the formal literature, including my own working paper (ssrn.com/abstract=6361077), the detection probability is a free parameter. Referees of that paper said the mechanism was sound and the institutional grounding thin. They were right, and this project attacks the grounding. The one mature comparison case is nuclear safeguards. The IAEA has run a treaty-based verification system over a dual-use industrial input for more than fifty years, with defined instruments, published detection goals, quantified inspection effort, and a public record of failures at Iraq, North Korea and Iran. Baker opened this comparison at treaty level (arxiv.org/abs/2304.04123) and Wasil et al. surveyed verification methods (arxiv.org/abs/2408.16074). Neither goes down to instrument-level performance, and that is the gap. Research questions. (1) Which safeguards instruments have functional analogues in compute governance, and where does the analogy fail? Material accountancy assumes a conserved, measurable stock; compute is a flow that can be split across providers and jurisdictions and recombined. We want a defensible account of which concepts survive that difference. (2) What does the published record support about detection performance, inspection effort and cost for each instrument, covering containment and surveillance, environmental sampling and complementary access? What is the corresponding evidence, or absence of evidence, for chip attestation, cloud provider reporting and aggregate energy signatures? (3) Given those parameters, which combinations of audit frequency, penalty schedule and third-party certification make truthful compute reporting incentive-compatible, and where does the required audit intensity exceed anything a regulator would fund? Workplan, twelve weeks. Weeks 1 to 3: shared orientation on safeguards implementation and compute verification; the team builds the coding scheme together (instrument, object verified, detection mechanism, documented performance, inspection effort, evasion history, compute analogue, transfer verdict, confidence) and pilots it on five instruments. Weeks 4 to 7: evidence extraction from IAEA safeguards implementation reports and technical documents, the arms control verification literature, and the compute governance corpus. Every cell has a citation and a confidence label; weekly written memos. Weeks 8 to 10: Jonas leads parameterisation, feeding extracted values into an existing compliance model and an agent-based stress test with noisy metering, evasion and collusion. The model and codebase exist and are handed over in week one, so nobody starts from a blank file. Weeks 11 to 12: writing, with a complete draft by Demo Day. Outputs. An openly licensed coded dataset and appendix, every claim sourced. A paper for a governance workshop or policy journal, with mentees as co-authors by contribution. A short note stating which verification architectures the evidence supports and which it does not. Why this team. I have a paper accepted for the IAEA Symposium on International Safeguards this November applying mechanism design to safeguards oversight, and I am a member of the Pugwash chemical and biological weapons working group, so the source material and the people who know it are accessible to us. Jonas Kgomo co-mentors and leads the computational side. Cohorts I supervised previously produced a workshop poster, a World Economic Forum article and extensions to the working paper this project builds on.
Theory of change
If governments adopt compute-based oversight, and the EU AI Act already does through its 10^25 FLOP threshold for general-purpose models with systemic risk, the design details decide whether the regime is enforceable or symbolic. Two failure modes follow from unevidenced enforcement assumptions. Regulators legislate thresholds nobody can check, which produces the appearance of control without the substance. And formal governance mechanisms are set aside by policy readers as ungrounded, so the analytical work never reaches the people writing rules. A sourced account of what inspection instruments detect, calibrated against the only large-scale precedent for verifying a dual-use industrial input, lets designers rule out architectures needing audit intensity no agency will fund, and lets them defend the ones that survive. That matters most for any future international arrangement over training compute, which is the pathway most likely to slow a race to unsafe capability. It also matters for domestic licensing now. The output is designed to be used rather than admired. A coded, openly licensed dataset with confidence labels is something other researchers can extend and contest, and something a regulator's staff can read in an afternoon. Related work: ssrn.com/abstract=6361077 and weforum.org/stories/2026/02/what-carbon-markets-can-teach-us-about-governing-frontier-ai/
Your role
Each mentee owns a named workstream and, by extension, a section of the paper. This is not assistance on someone else's project. Two roles. Evidence leads (two mentees) each take a set of inspection instruments, apply the coding scheme, source every claim, and write the corresponding section. They decide the transfer verdicts and defend them against challenge from the rest of us. Simulation lead (one mentee) takes the existing model and codebase, implements the stress tests, and writes the results section. A fourth mentee, if the applicant pool is strong, extends the evidence side to non-nuclear inspection regimes such as chemical weapons or financial supervision. Autonomy is high on method within a workstream and low on scope. We fix the coding scheme jointly in week three and then keep to it, because the value of the output depends on consistency across rows. Judgement calls stay with the mentee who made them, recorded in the confidence field rather than overruled by me. Authorship is by contribution and discussed openly in week two, not decided at the end. Mentees whose work supports a section are co-authors on it. Several people from previous cohorts continued working with me after the programme closed.
Prerequisites
One of the two profiles below is enough. You do not need both. Evidence track. You can read technical or legal source material at length and pull structured claims out of it. You have written at least one substantial researched piece where you had to weigh conflicting sources and say which you believed. Comfort with treaty text, inspection protocols or regulatory documents helps. No prior nuclear knowledge is needed; the reading list is provided and the first three weeks are shared orientation. Simulation track. You can read and extend someone else's Python codebase without hand-holding, and you have taken microeconomic theory, game theory or mechanism design at an advanced undergraduate level or beyond. You should be able to state what a Nash equilibrium of a simple audit game is and compute one. Across both: you write well in English, you finish what you start, and you can commit ten hours a week for twelve weeks. We would rather have three people at ten reliable hours than five at fifteen intended ones.
Location preference
No geographical requirement. One weekly team call, fixed in week one within 13:00 to 17:00 Central European Time, which is workable from the Americas and from Asia at the edges. Everything else is asynchronous, in writing.
Application question(s)
Answer question 1 or 2 depending on which track you are applying for. Answering both is welcome and not expected. These should take twenty to thirty minutes.
- The IAEA defines a significant quantity of plutonium as 8 kg and sets a detection timeliness goal of one month for unirradiated direct-use material. Propose the closest equivalents for a compute permit regime. What is the significant quantity of compute, and what is the timeliness goal? Say what part of the analogy fails and why. (300 words)
- A regulator audits 5 percent of declared training runs each year and fines a detected under-reporter three times the value of the permits evaded. Assume risk neutrality and no reputational cost. Is truthful reporting an equilibrium? Show the calculation. Then name the assumption in that setup you distrust most, and say what evidence would change your mind. (250 words)
- Link one piece of analytical writing you produced, ideally from research. In two sentences, say what you would change about it now.
About the mentors

Joel Christoph works on the economics and institutional design of compute governance, meaning the question of how limits on frontier training runs can be enforced when a regulator cannot observe everything a developer does. He is a Research Associate and Project Coordinator on the Graduate Programme on Existential Risks to Humanity at FernUniversität in Hagen, developed with the Centre for the Study of Existential Risk network. His working paper on risk-weighted compute permit markets under imperfect monitoring is on SSRN, and his writing on compute and AI regulation has appeared in Lawfare, Tech Policy Press and with the World Economic Forum. He was a Summer Fellow at the Centre for the Governance of AI, a PIBBSS fellow, and a Technology and Human Rights Fellow at Harvard Kennedy School's Carr-Ryan Center, and has an MRes in Economics from the European University Institute.
He also works on nuclear arms control, with a paper accepted for the IAEA Symposium on International Safeguards in Vienna in November 2026 and membership of the Pugwash chemical and biological weapons working group. This project draws on both halves of that record. Joel has mentored three previous SPAR rounds and led project teams at AI Safety Camp, whose members contributed extensions to the compute permits working paper and co-authored conference posters. Jonas Kgomo co-mentors this project; he co-founded the Equiano Institute with Joel, co-leads a compute governance research cohort with him, and will lead the computational workstream. We run a weekly team call with written feedback in between, and we aim for a public, co-authored output from every round.

Jonas Kgomo is Founder of the Equiano Institute. His work focuses on responsible AI and digital infrastructure governance, with particular attention to how institutions in emerging markets can adopt AI safely and credibly. He has experience with multi-stakeholder policy processes and applied governance research, including work on responsible AI in Kenya. He previously co-mentored a Spring 2026 SPAR project on market mechanisms to incentivize responsible AI development with Joel Christoph.