Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Mitigating Intentional Loss of Control Risk Through Interoperability Standards for Agentic AI

AI control International governance Technical governance

Within a few years, when self-replication is plausibly within reach of open-weight systems, the binding constraint on an agent deployed to operate without a controlling human principal will likely not be the model, but whether the rest of the agent economy will discover, authorize, transact with, or pay it. This project maps who actually holds change control over the agent identity and trust layer, evaluates the competing architectures against an explicit intentional-loss-of-control threat model, and feeds the result into live standards and Geneva policy discussions before network effects settle the question.

About the project

The problem. Standards are where technical detail quietly becomes de facto international law. The Internet is the cautionary case: IP won the protocol wars on network effects, and its identity and trust layers were bolted on afterwards. Thirty years later BGP still runs on trust, the web PKI still depends on every certificate authority in the store behaving, and the neutrality of the DNS root still turns on who holds jurisdiction over ICANN. None of it has been fixable, because change control was settled before anyone understood what was at stake.

Agent-to-agent interoperability is at that stage now, and moving faster. MCP moved under a Linux Foundation body in late 2025; its authorization model is conventional OAuth and the specification has no concept of agent identity. A2A reached its first stable version under the Linux Foundation, where an agent identifies itself through a card it publishes about itself and signing that card is optional. Agent payments are moving fastest of all, with Google's Agent Payments Protocol donated to the FIDO Alliance in 2026 — and it is the partial counterexample worth studying, because it anchors trust in cryptographically signed user mandates rather than in agent self-declaration. At the IETF, agent identity work remains a scatter of individual drafts, none adopted by a working group; a proposal to open a working group on agent delegation and traceability was declined in July 2026. ITU-T established a focus group on trust and identity for agentic AI only in July 2026, with first work items targeted for 2027 and beyond. ISO/IEC has no agent-specific standard in development.

The safety argument. Intentional loss of control is the case where a human deliberately deploys an agent to operate without an accountable controlling principal — designed to resist shutdown, to re-establish itself elsewhere, or to replicate. As capability diffuses through open weights, mitigation at the model layer weakens: you cannot rely on a developer's safeguards for a system whose weights are public and whose fine-tuning is cheap. What remains are the downstream chokepoints, and in an economy where agents transact with agents, the chokepoint is admission to the network.

The architecture decides whether that chokepoint exists at all. Under self-declared identity, an adversarial deployer forges or clones a credential, re-registers after any block, and pays no cost; revocation has nothing to bite on. Under issued, revocable identity with an auditable delegation chain back to an accountable legal person, revocation becomes a real intervention at the transaction layer, unattributed agents can be refused by counterparties and payment rails, liability attaches to someone, and incidents can be reconstructed.

That is the claim the project tests: organized, issued agent identity is a control lever; self-declared identity is not. It is not a free lever, and the project has to be honest about that or it will not survive contact with the standards community. The same identifier that lets you revoke a replicating agent lets a state bind every agent action to a real-world human identity. A registry is a governance object with an owner and a jurisdiction, which is the ICANN question again. The research question is therefore not whether agent identity is good, but: which architecture buys the most loss-of-control mitigation per unit of surveillance capability, centralization, and foreclosure risk — and who has to hold change control for that to remain true?

Research questions. (1) Who holds decision rights over the agent identity and trust layer today, by body and by decision date, and where does that move over the next 18 months? (2) What does an intentional-loss-of-control threat model demand of an identity architecture, derived from concrete scenarios rather than enterprise IAM best practice? (3) How do the live design families score against those requirements and against privacy, centralization, jurisdictional exposure, and adoption feasibility? (4) Which properties must live in the protocol and which can be left to policy or procurement — following Collingridge, what is cheap to change now and prohibitively expensive later? (5) What is the minimum viable ask a safety-motivated actor could bring into a named process in the next twelve months, and to whom?

Work packages. WP1 (weeks 1–3): change-control map across the protocol foundations, internet standards bodies, payments alliances, identity community, and intergovernmental track — owner, stage, change-control mechanism, funding, next decision date. WP2 (weeks 3–6): three to four intentional-loss-of-control scenarios walked through the transaction lifecycle from discovery to revocation, yielding a requirements set. WP3 (weeks 5–9): comparative scoring of the design families against those requirements and the counterpart criteria. WP4 (weeks 9–12): intervention analysis — which ask, which body, which deadline — plus one brief for standards participants and one for diplomats.

Deliverables. A change-control map with a decision calendar; a requirements set; a scored comparative assessment; a policy brief of roughly 3'000 words plus a shorter technical-community version. Stretch goal: a submitted comment to a named open process.

Method. Document analysis of primary standards artifacts, all public. Eight to fifteen structured expert interviews, using my network for access. Scenario-based reasoning rather than quantitative modeling. Comparative scoring against published criteria with explicit uncertainty, so a reader who disagrees with a weight can recompute rather than argue. No compute required.

Theory of change

While the biggest loss of control risk currently is internal deployment in frontier AI labs, I expect AIs with high autonomy risk to diffuse. We cannot rely on a developer's controls for a system whose weights are public and whose fine-tuning is cheap. The mitigations that survive if capability diffuses are downstream of the model, where an agent has to interact with the world. In an economy where agents transact with agents, the decisive one is admission to the network — whether counterparties, service providers, and payment rails will deal with an agent that cannot produce a valid, revocable identity traceable to an accountable principal.

Whether that chokepoint exists is being decided now, in protocol specifications, by people who are solving an enterprise access-management problem rather than a catastrophic-risk problem. Self-declared identity gives an adversarial deployer a free re-registration after every block. Issued and revocable identity makes revocation a real intervention, lets platforms refuse unattributed agents, and attaches liability to a person. This is a rare case where a safety property can be built into infrastructure rather than negotiated between states afterwards — and a rare case where the window is identifiable and short, because network effects will soon settle these questions permanently.

The output is therefore not a preprint but a mapped set of asks: which change to which document in which body, by when. I have done the analogous analysis before: at ETH Zurich's Center for Security Studies I published a study of the politics of Internet protocol design covering New IP, SCION, ICANN's jurisdiction problem, and the IETF–ITU conflict. I am now based in Geneva launching a new AI governance institute, and recently convened a private roundtable on AI standards with major AI labs, global standards bodies, a leading payment network, and three national AI safety institutes, which generated follow-up interest this project is designed to feed.

Your role

High autonomy. Each mentee owns work packages end to end and is the author of what they produce; I set direction, review drafts, open doors, and argue with conclusions. This is not a project where I hand over a task list.

With two mentees the split is: one owns the change-control map and the comparative assessment (WP1, WP3), the other owns the threat model and the intervention analysis (WP2, WP4), with a shared weekly synthesis where each has to defend their work to the other. Both tracks feed a joint output, so mentees will interact substantially rather than working in parallel silos.

Concretely, a mentee will: read primary standards artifacts closely and summarize what they actually commit implementers to; build and maintain a structured landscape database; write and run expert interviews, including drafting the outreach; construct scenarios and stress-test candidate architectures against them; and draft sections of the final brief under their own name.

What I provide: weekly direction and prioritization, written feedback on drafts, introductions from my network for the interview program, and the route by which the output reaches decision-makers. What I do not provide: line-by-line technical supervision of protocol analysis. I would likely put uncertain technical readings in front of trusted experts in my network rather than adjudicating them myself.

Published outputs carry mentee authorship. If the work is good, I will take it into Geneva rooms with the mentee's name on it.

Prerequisites

Able to read a technical specification or standards draft closely and explain what it actually obliges implementers to do — without needing to implement it. If you have never opened an IETF or W3C draft, that is fine; if the idea of reading one carefully for two hours sounds unpleasant, this is the wrong project. Comfortable writing clear analytical prose for a policy audience. Most of the output is written. Able to work independently on an open-ended question and come back with a structure rather than a list of links. Willing to email strangers, request interviews, and run them. Some grounding in AI safety arguments, specifically why loss of control is a concern at all. You do not need to agree with the framing — informed skepticism is useful here — but you should be able to engage with it.

Useful but not required: background in internet governance, identity and access management, security engineering, standards participation, or law. Experience with a slow institutional process of any kind, including as a participant.

Explicitly not required: machine learning research experience, ability to train models, publications, or a technical degree.

Application question(s)

Which forum is likely to be most influential in setting agentic AI standards and why? (200 words)

Pick any agentic AI protocol or draft specification of your choice. Read it, then highlight some of its (potential) political implications (e.g., interoperability, change control, privacy, ability to shut an agent down, and intentional loss of control). (350 words)

Link to a writing sample, ideally highlighting analytical or research writing. (No word limit.)

About the mentor

Kevin Kohler

Kevin Kohler

Simon Institute, UN University, AGI Preparedness Institute

View profile

Hi! I'm Kevin, I'm based in Geneva, and I work on AI governance and institutional design questions.

From 2019 to 2022 I was a researcher at the Center for Security Studies at ETH Zurich, where I produced a report on the politics of Internet standards and future Internet architectures, which is the intellectual ancestor of this specific project proposal.

Since then I have worked mostly on the multilateral side of AI governance: I engaged the UN General Assembly negotiations that established the first UN institutions dedicated to AI, served on the writing and editing support team for the UN Scientific Panel on AI's first report, and have run various AI governance briefings for diplomats in Geneva, including a private roundtable on AI standards with major AI labs, standards bodies, and safety institutes.

I am currently leading research and multilateral projects at the Simon Institute, though I'm in the process of launching a new institute on AGI preparedness. I also blog on AGI futures and institutional design at machinocene.com.

I have mentored across three cohorts of the Cambridge ERA:AI fellowship. As a mentee, you should be in the driver's seat, but I can be useful for working out which questions are policy relevant, how to write something a decision-maker will read, and for opening doors when you need interviews. I should also flag I am a policy researcher, not a standards engineer, so on hard technical questions I will connect you with someone who can settle them. Expect a weekly call, written feedback on demand, and your name on what you produce.

Similar projects