Fall 2026 mentee applications are open! Apply to research projects by August 18. Apply now

All Fall 2026 projects

Global AI Risk Observatory

AI strategy Societal impacts National policy

The Global AI Risk Observatory analyses corporate disclosures at scale (~1M documents: annual reports, earnings calls, investor presentations) using LLM auto-graders to track how companies worldwide report AI adoption, risks, and dependencies. Building on a UK AISI-funded pilot of 9,821 UK annual reports, we're expanding to global coverage and translating disclosure trends into policy-relevant findings for AI governance and societal resilience.

About the project

The pilot found that AI risk mentions in UK annual reports grew roughly fourteen-fold between 2020 and 2025 (3.2% → 43.2% of reports), yet only 4.2% of reports contained substantive disclosure describing specific mechanisms or controls — a widening gap between disclosure volume and quality with direct implications for regulators.

Theory of change

Societal resilience to transformative AI requires knowing where AI dependencies and risks are accumulating across the economy before failures materialise — and today, no systematic visibility exists. Corporate disclosures are the only legally mandated, continuously updated record of how companies understand their own AI exposure. Our UK AISI-funded pilot (riskobservatory.ai) showed LLM pipelines can extract validated signals from them at scale, and surfaced actionable blind spots: critical infrastructure sectors like Energy and Data Infrastructure disclose AI risk at the lowest rates, and only 4.2% of disclosures are substantive. Scaling this globally creates monitoring infrastructure for governments: better measurement of AI dependency → better-targeted governance → greater resilience through the transition to transformative AI

Your role

Two complementary roles. The technical fellow will own extensions to the LLM classification pipeline: adapting prompts and taxonomies to new jurisdictions and document types, building validation sets, running large-scale classification jobs, and producing statistical analyses. They'll have high autonomy over implementation choices within an established architecture, with the mentor reviewing design decisions and validation methodology. The governance fellow will own the translation layer: mapping disclosure patterns to regulatory contexts across jurisdictions, cross-referencing against external datasets, and drafting policy-facing analyses. They'll shape which questions we prioritise — the agenda is deliberately exploratory, and we expect fellows to propose and pursue directions the data suggests. Both fellows will be credited contributors on resulting publications, and the two roles are designed to collaborate: the governance fellow's questions drive what the technical fellow measures, and vice versa.

Prerequisites

Technical fellow (LLM data-analysis pipeline):

Highly proficient in Python Has built at least one non-trivial LLM application using model APIs (e.g. classification, extraction, or evaluation pipeline — personal projects are fine) Comfortable with prompt engineering and structured outputs (e.g. JSON mode, schema-constrained generation) Basic applied statistics: can compute and interpret agreement metrics and sanity-check results at scale

Governance & policy fellow:

Strong analytical writing, demonstrated by a writing sample Working familiarity with at least one AI regulatory regime (e.g. EU AI Act, UK framework, US SEC disclosure rules) Comfortable reading quantitative results and reasoning carefully from data to policy implications No coding required, but must be willing to work hands-on with data outputs (spreadsheets, dashboards)

Application question(s)

Q1 (Technical fellow applicants): Our pipeline classifies passages from annual reports as AI adoption, risk, vendor, or harm mentions. In validation, our LLM classifier often assigned more labels per passage than the human annotator did, and manual inspection suggested many of the extra labels were valid. Does this mean the classifier is better than the human, and how would you design a validation process that answers that question rigorously rather than by assumption? What would you report as your headline accuracy metric, and why? (300 words)

Q2 (Governance fellow applicants): In our UK data, 43.2% of 2025 annual reports mention AI as a risk, but only 4.2% do so substantively (naming specific mechanisms, controls, or targets). We can think of at least three explanations: annual reports structurally favour hedged language; AI governance is too immature for boards to say anything specific; or companies strategically choose vagueness to limit liability. Pick the explanation you find most plausible, defend it, and state what evidence, from our data or elsewhere, would change your mind. Each explanation implies a different regulatory response: what does yours imply? (300 words)

Q3 (All applicants): Link to one relevant work sample: for the technical role, a repository or write-up of an LLM project you built; for the governance role, an analytical writing sample, ideally empirical or policy-focused.

About the mentor

Bart Jaworski

Bart Jaworski

MATS

View profile

I'm currently at MATS working with Victoria Krakovna on realistic model organisms of scheming.

In the past I started and run a a few companies (some of which even made money) and most contracted UK AISI to build the AI Risk Observatory (riskobservatory.ai).

Similar projects