Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Auditing Frontier AI Compliance via an EU Code of Practice Tracker

EU policy Technical governance Lab governance

The EU AI Act's Code of Practice sets out what frontier AI providers must disclose about how they manage and mitigate risk. This project builds a tracker that measures how well each provider actually complies with the CoP. The hard part of this project is going to be measurement: turning vague legal obligations into atomic, checkable indicators that measure compliance. The work in the project involves designing indicators, doing literature reviews of model evaluations to define what a complete disclosure looks like, human review and sign-off on every indicator and score for every provider assessed and evaluating a multi-model AI pipeline that does the scoring. We need a few different technical profiles.

About the project

Full enforcement of the EU AI Act for general-purpose AI starts on 2 August 2026, and the Code of Practice is the detailed rulebook for what providers must do and disclose. Checking that across many providers and many obligations by hand doesn't scale, so we've built a multi-agent DELPHI pipeline that does a first pass automatically. It reads the Code text and turns each obligation into a set of scoring indicators, using a dynamic multi turn debate setup. Panels of models propose indicators, adversarially red-team them, and arbitrate the disagreements. It then scores each provider's public documents against those indicators, clause by clause, tying every judgment to a quoted piece of evidence and attaching a confidence estimate. The pipeline gives us breadth and speed. What it can't give us is the judgment to know when an indicator is a weak, incomplete, or gameable reading of an obligation.

I need collaborators on this project, who can work on several different sub-parts:

  • You can help assess if the indicators actually capture what the obligation requires. Is it atomic and checkable, or vague? Could a provider satisfy it on paper while doing nothing meaningful — is it gameable? Where an indicator is weak you improve it; where the pipeline missed something, or produced something a lab could safety-wash its way past, you design a new indicator that closes the gap and resists gaming. The role is less about generating indicators from a blank page and more about being the critical expert layer on top of the pipeline that makes its output trustworthy.

  • It is useful to have someone who is familiar with evaluations, risk modeling, and is willing to read through system cards and safety policies for new models. Evaluation literature reviews define what a complete disclosure actually looks like: what belongs in a real risk model, a real model or system card, or the state-of-the-art evaluations for a given risk domain.

  • You can help build the multi-agent pipeline that I described in the first paragraph. In this case, you would be doing a lot of coding and engineering work. Additionally, an obvious research question sits underneath the whole project: how much can you trust a pipeline like this? That's the meta-evaluation part of the project that we would want to turn into a research paper. Does a panel drawn from several model vendors produce better indicators than one vendor alone? Does a bigger panel or a longer red-team round help, or just cost more? Are some models better at proposing indicators and others at judging them? Do different providers have different DCG (generation-vs-discrimination/critique) gaps? A hands-on critique of the pipeline's output is the ground truth that makes these questions answerable. You can see where it fails, and measure whether a design change fixes it.

Alongside this there's scoring and coverage work: re-scoring providers across multiple runs to measure run-to-run variance and put confidence on the numbers, and extending coverage to more providers, including companies from China.

Theory of change

The EU AI Act's rules for general-purpose AI are one of the few binding levers on frontier labs, but they are only useful if someone can independently show whether providers meet their commitments. A public, evidence-backed compliance dataset, landing as enforcement begins (Aug 2nd 2026), creates that pressure and gives the AI Office, journalists, and civil society a shared set of facts to argue from.

Beyond having public transparency for regulatory compliance levels, we also need to move towards data backed AI governance, and model assisted workflows. Many benchmarks and evaluations already come paired with LLM-as-a-judge protocols. Beyond benchmarking, we can use similar techniques to scale our efforts to govern AI: grading safety cases, checking compliance, scoring frameworks. If we don't know when those systems are trustworthy, and when a panel of models just launders one model's bias at higher cost, we'll build governance on sand. This project attacks that question on a real regulation, and trains someone who can build measurement pipelines a policymaker can rely on when making judgements.

Your role

We are already a team of 3 that you will be joining. Review and scoring is shared across the team. I expect mentees to be autonomous and self directed. I will be coordinating the team, and will provide weekly guidance and mentorship.

Prerequisites

We're looking for one of the following profiles. A strong generalist can span multiple.

  • Quantitative social science, or a similar background in turning qualitative material into valid quantitative indicators. This is measurement work: construct validity (measuring the thing you claim to, not a proxy for your own opinion), decomposition (breaking a qualitative obligation into atomic, comparable, checkable units), and aggregation (how those units combine without one failure hiding behind many passes). Content analysis, coding schemes, psychometrics, survey methodology, or empirical social science all count.
  • An engineer fluent with LLM tooling, or who has worked with multi-agent workflows before, for the pipeline and the meta-evaluation study. Strong Python, comfortable in a real research codebase, and able to design and run a clean experiment on the pipeline.

Across all: able to read regulatory and technical text closely and precisely; genuinely care about the line between compliance and safety, and about construct validity; and self-directed enough to take a scoped brief and return a defensible result without hand-holding.

Application question(s)

  1. A measure in the Code has several requirements. There are two ways to turn them into a numeric score. For a compliance tracker that has to be reproducible and defensible to a hostile reader, which do you choose? Why? and what does your choice give up? (≤400 words):

(a) Importance weighting: give each indicator a weight for how important that requirement is (this one 0.4, that one 0.15) and take the weighted average. This lets you say some requirements matter more than others. The weights are many interdependent judgment calls that must stay consistent every time an indicator is reworded or re-split, which is a heavy maintenance cost; a weighted average still lets several easy passes buy back one critical failure (heavier weights only change the exchange rate, they don't stop the trade), and lastly weighting can be seen as encoding subjective opinions. (b) Clause-mass weighting: every clause counts equally; the measure's score is a flat average over its clauses, ignoring how they're grouped into indicators, so splitting or merging indicators can't move the score. Measures are themselves combined equally at the level above, so the number of clauses a measure has changes how much each clause counts toward the commitment. If a commitment has two equally-weighted measures, a clause in the 4-clause measure counts ⅛ toward the commitment while a clause in the 40-clause measure counts only 1/80.

  1. If you are interested in working on the engineering side of the project, then please also answer the following question: Design an experiment to test whether one panel-design choice (vendor mix, panel size, red-team length) improves the indicators or scores. Give your metric, your control, and the biggest problem. (≤300 words)

About the mentors

Markov Grey

Markov Grey

CeSIA (French Center for AI Safety)

View profile

I am the head of technical AI governance at the French Center for AI Safety (CeSIA). Right now I am designing harmful manipulation evaluation standards for the European Commission's AI Office. I have also been the co-founder and CTO of Equilibria Network, focusing on collective intelligence and complex systems safety. Before this, I was researching measurement standards for AI safety, risk modeling and was leading writing for the AI Safety Atlas textbook. I have previously been a scriptwriter for Rational Animations, and trained in cybersecurity, mathematics, and computer science.

Charbel-Raphael Segerie

Charbel-Raphael Segerie

CeSIA - Centre pour la Sécurité de l'IA (French Center for AI Safety)

View profile

Charbel-Raphael Segerie has extensive experience in AI safety field-building, education, and content creation. He was previously Head of AI at EffiSciences, funded ML4Good, was CTO of a startup, and worked in different French research institutions (Inria, Neurospin). He is an OECD AI expert. He teaches AI safety in ENS, which is one of the only university accredited courses in the EU on AGI safety. His research focuses on identifying emerging risks in artificial intelligence, improving current safety methods such as RLHF and interpretability, and advancing safe-by-design AI approaches. Additionally, he contributed to AI evaluation efforts and collaborated on the EU AI Office's Code of Practice for general-purpose AI systems.

Similar projects