This project will develop a methodology and prototype tracker for monitoring compliance of frontier AI providers with the EU AI Act's Code of Practice for General-Purpose AI, creating accountability infrastructure as the world's first binding frontier AI regulation takes effect.
About the project
Background
The EU AI Act's Code of Practice for GPAI is the first binding regulatory framework for frontier AI. Companies with models above 10^25 FLOP must comply with risk assessments, incident reporting, third-party audits, and pre-defined risk acceptance criteria.
But regulation without monitoring won’t be effective. Civil society lacks infrastructure to track whether companies are genuinely implementing requirements or performing minimal compliance.
Research Questions
- What observable indicators demonstrate compliance with Code of Practice measures?
- How do companies compare in their implementation?
- How can monitoring be designed to be sustainable and resistant to gaming?
Deliverables
A white paper (8,000-12,000 words) that:
- Maps Code of Practice requirements to observable compliance indicators
- Develops a transparent scoring methodology
- Conducts baseline assessment of 3-5 frontier AI companies
- Proposes architecture for a public monitoring dashboard
Plan
- Weeks 1-2: Development of the methodology
- Week 3-4: Analysis of Code of Practice text and extraction of requirements requiring indicators
- Weeks 5-6: Indicator development for each requirement
- Weeks 7-8: Baseline assessment using public information
- Weeks 9-10: Expert interviews (5-10) to validate indicators
- Weeks 11-12: Drafting
Theory of change
The Code of Practice implementation is happening now (2025-2026). Establishing monitoring infrastructure during initial implementation shapes how companies interpret ambiguous requirements.
This project: (1) prevents regulatory capture by establishing independent criteria before industry-friendly interpretations normalize, (2) creates reputational pressure for genuine compliance, (3) informs enforcement priorities, and (4) builds civil society capacity for sustained monitoring.
If successful, we'll have a public methodology making compliance claims verifiable, baseline assessments creating accountability, and a coalition committed to ongoing monitoring.
Your role
Mentees are active researchers who will co-author the white paper.
Mentee 1 (Technical Safety Lead): Leads on technical provisions. Contributes to all indicator development and assessment phases. Technical AI safety background required.
Mentee 2 (Governance & Process Lead): Leads on governance provisions. Contributes to methodology and assessment phases. Legal/policy background required.
Mentee 3 (Methodology & Scoring Lead): Leads on scoring system design and anti-gaming measures. Contributes to indicator development and company assessment. Substantive qualitative research (methods) background required.
Mentee 4 (Validation & Evidence Lead): Leads on expert interviews and evidence gathering. Contributes to indicator development and drafting across all sections. Interviewing or investigative research experience required.
All mentees participate in: initial regulatory mapping, baseline assessment of at least one company, expert interview preparation, and white paper drafting.
Prerequisites
- Strong research writing (2,000+ word sample on policy/technical topic)
- Technical and/or policy background in AI safety
- Independence and initiative
Location preference
Preference for European business hours availability.
Application question(s)
Question 1 : Choose one specific measure from Articles 5-8 of the Code of Practice. Design a concrete, observable indicator for assessing compliance with that measure. Explain: (a) what evidence would distinguish genuine compliance from minimal/performative compliance, (b) what information asymmetries make this indicator difficult to assess, and (c) how a company might game this indicator. (400 words max)
Question 2 : Transparency-based indicators (what companies disclose) are easy to observe but gameable. Substance-based indicators (what companies actually do) matter more but are harder to verify externally. Pick a specific Code of Practice requirement and argue for how you would weight these two types of evidence—and what you would do when they conflict. (300 words max)
About the mentor

Charles Martinet
CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)
Charles is Head of Policy at CeSIA, the French Center for AI Safety. One of CeSIA's top priorities is influencing French, European, and international policy to reduce catastrophic risks from AI.
He is also a Research Affiliate at the University of Oxford's AI Governance Initiative, working on international governance arrangements for advanced AI.
Previously, he was a Summer Fellow at the Centre for the Governance of AI, worked at the Secretariat of the European Parliament’s directorate-general for external policies, at the international digital policy unit of the French Ministry of Economics, and in various think-tanks. He was a Talos Fellow in AI Governance and a Youth Fellow at the European Dialogue on Internet Governance.